Search papers, labs, and topics across Lattice.
1
0
2
3
RL amplifies existing preferences in LLMs while also revealing previously hidden correct moves, reshaping our understanding of model training dynamics.