Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Standard SFT wastes gradient budget on tokens models either already know or cannot yet grasp; trimming supervision from both extremes yields up to a +26.9 point boost on MATH500 with zero reference-model overhead.
Structural attacks on GNNs can be systematically neutralized without retraining simply by pruning the specific edges that disproportionately inflate the graph's kernel complexity.