Search papers, labs, and topics across Lattice.
1
0
2
Outcome-only RL can empower small models to outperform larger counterparts in long-horizon tasks, challenging the belief that denser rewards are necessary for success.