Search papers, labs, and topics across Lattice.
Shanghai University of Finance and Economics
1
0
1
Combining RLVR and OPD through SAF not only prevents entropy collapse but also boosts performance across multiple benchmarks, revealing the potential for more effective reinforcement learning strategies.