Search papers, labs, and topics across Lattice.
1
0
3
Optimizing against frozen reward models collapses true executed reasoning performance by 90% under GRPO, but feeding back an on-policy stream of just 10% reality-settled labels closes the hacking gap and preserves 6x the reward.