Search papers, labs, and topics across Lattice.
This review critically examines On-Policy Self-Distillation (OPSD) in language models, highlighting its potential to reduce the reliance on a larger teacher model by using the model itself as the teacher with privileged information. It identifies a significant failure mode called "collapse," where the model progressively narrows its reasoning paths due to biases in the self-distillation process. By analyzing the factors contributing to collapse鈥攕ignal application, the nature of privileged information, and the timing of signal changes鈥攖his work provides a structured framework for understanding and addressing these challenges in mathematical reasoning tasks.
Collapse in On-Policy Self-Distillation narrows reasoning paths, revealing critical biases that could undermine model performance.
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were promising, with accuracy comparable to reinforcement learning at a fraction of the generated tokens. But the same asymmetry that produces the signal also biases it. One failure mode now dominates the field: collapse, the progressive narrowing of the set of reasoning paths the model can produce. Collapse is not specific to OPSD, though privileged information aggravates it. This review treats collapse as a symptom governed by three levers: (i) where the signal is applied, that is, how tokens are weighted; (ii) what the teacher is shown, that is, the nature of the privileged information; and (iii) when the signal changes, that is, the teacher's dynamics and the decay of guidance. We restrict our scope to mathematical reasoning, where the method originated and where its failure modes are best documented. We report no new experiments. The contribution is structural: a shared vocabulary for phenomena named differently across papers, and a clear line between what is settled and what is still disputed.