Search papers, labs, and topics across Lattice.
The authors formulate on-policy distillation (OPD) as an explicit probability transport problem鈥攖ermed RouteOPD鈥攖hat directly couples student-excess source tokens with teacher-deficit destination tokens. This addresses a fundamental flaw in sampled OPD objectives, which reduce rich teacher distributions to scalar per-token credits and cause uncontrolled background probability leakage. Evaluating across four teacher-student configurations and four mathematical reasoning benchmarks, RouteOPD consistently outperforms standard sampled reverse-KL OPD with higher routing fidelity and tighter probability redistribution.
Scalar rewards in on-policy distillation tell tokens to gain or lose probability without specifying where that mass should actually go鈥攆raming distillation as explicit pairwise probability transport eliminates this background leakage and beats reverse-KL across reasoning tasks.
On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. We recast OPD as teacher-guided probability transport and propose RouteOPD (Routed On-Policy Distillation), which decomposes local teacher--student disagreement into student-excess sources and teacher-deficit destinations and couples them into explicit transport pairs. RouteOPD optimizes pairwise log-odds toward jointly realizable targets obtained from a bounded teacher potential, while adapting the transport budget to the concentration of teacher demand. This formulation directs updates toward teacher-preferred destinations and controls their magnitude within a single transport operator. Experiments across four teacher--student settings and four mathematical-reasoning benchmarks demonstrate that RouteOPD consistently outperforms sampled reverse-KL OPD, with improvements accompanied by higher routing fidelity and lower background leakage. These results demonstrate the effectiveness of explicitly modeling probability transport in on-policy distillation.