Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.
Mixture-of-Experts models might be hiding more of their reasoning than we thought, thanks to a newly quantified "opaque serial depth" metric.