Search papers, labs, and topics across Lattice.
1
0
2
Standard RLVR observations are mathematically insufficient to isolate or fix verifier exploitation, proving that reasoning models cannot unlearn rewarded errors without external audit signals.