Search papers, labs, and topics across Lattice.
2
0
3
6
Closing 97% of the performance gap between student and teacher models while cutting rollout steps by nearly three times reveals the power of adaptive domain scheduling in multi-teacher distillation.
A mere 0.01% of tokens can destabilize LLM reinforcement learning, but masking their gradient updates unlocks significant performance gains.