Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.
Fine-tuning LLMs with RL can boost self-explanation faithfulness from nearly zero to 0.664, revealing surprising cross-intervention generalization capabilities.