Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.