Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
On-policy LLM distillation does not actually need precise advantage magnitudes: retaining merely the directional sign of token advantages matches standard distillation performance while Total Variation smoothing eliminates late-stage training instability.
FastMTP speculative decoding doubles the decoding speed on low-budget GPUs while maintaining high accuracy in document parsing.