Search papers, labs, and topics across Lattice.
This study critically examines On-Policy Self-Distillation (OPSD) by introducing a novel approach called OP虏SD, which utilizes problem-solution pairs from different examples instead of relying on privileged information from the same problem. The findings reveal that the performance improvements attributed to OPSD are not solely due to access to the reference solution but are significantly influenced by the contextual behavior of the teacher. By validating this through experiments across three models and mathematics benchmarks, the research highlights the importance of context in the distillation process, suggesting a shift in how OPSD is understood and applied.
OPSD gains stem more from the teacher's contextual influence than from privileged access to problem-specific solutions.
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with $\mathrm{OP}^{2}\mathrm{SD}$ (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, $\mathrm{OP}^{2}\mathrm{SD}$ improves over the base model, remains competitive with OPSD. The success of $\mathrm{OP}^{2}\mathrm{SD}$ implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.