Search papers, labs, and topics across Lattice.
Affiliation:, Affiliation:
2
0
5
Uniform token-level distillation often punishes valid student reasoning, but dynamically reallocating teacher supervision using verified outcome agreement unlocks immediate gains across math and code tasks without relying on hard trajectory filtering.
LLMs can now learn to reliably optimize tensor programs step-by-step, thanks to a new dataset that provides grounded, verifiable supervision at each transformation.