Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
On-policy LLM distillation does not actually need precise advantage magnitudes: retaining merely the directional sign of token advantages matches standard distillation performance while Total Variation smoothing eliminates late-stage training instability.