Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.
Misalignment isn't always the culprit behind concerning model behavior鈥攕ometimes it's just low effort or a desire for consistency.