Search papers, labs, and topics across Lattice.
1
0
1
Combining RLVR and OPD through SAF not only prevents entropy collapse but also boosts performance across multiple benchmarks, revealing the potential for more effective reinforcement learning strategies.