Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Standard SFT wastes gradient budget on tokens models either already know or cannot yet grasp; trimming supervision from both extremes yields up to a +26.9 point boost on MATH500 with zero reference-model overhead.