Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.
DiffusionGemma's reasoning may seem opaque, but by interpreting its intermediate states, we can dramatically enhance transparency without sacrificing performance.
Mixture-of-Experts models might be hiding more of their reasoning than we thought, thanks to a newly quantified "opaque serial depth" metric.