Search papers, labs, and topics across Lattice.
This paper investigates optimal self-distillation (SD) for rectified flow (RF) by analyzing the impact of a suboptimal teacher velocity field on a student model trained with a mixture of true and teacher velocities. The authors derive a closed-form expression for the optimal mixing coefficient, demonstrating that positive mixing can correct under-regularized teachers while negative mixing addresses over-regularized ones, leading to improved velocity risk estimates. Experimental results across various models indicate that optimal self-distillation enhances mode recovery and reduces generation error compared to both the teacher model and conventional distillation methods.
Optimal self-distillation can significantly enhance generative model performance by intelligently mixing teacher and true velocity signals, correcting both under- and over-regularization issues.
Modern generative models are increasingly trained using model-generated signals, creating both opportunities for self-improvement and risks of collapse. We study optimal self-distillation (SD) for rectified flow (RF): given a suboptimal teacher velocity field, can a student trained on a mixture of true RF velocities and teacher velocities provably improve the teacher? For linear RF with ridge regularization on fixed interpolation pairs, we prove an exact affine path identity, derive the optimal mixing coefficient in closed form, and show strict improvement in integrated velocity risk whenever the teacher risk is nonstationary along the regularization path. The optimal coefficient obeys a sign rule: positive mixing corrects under-regularized teachers, while negative mixing corrects over-regularized teachers. We also give one-shot generalized cross-validation (GCV) and validation tuning procedure that avoids grid search over mixing weights and repeated refitting. Combining this theorem with RF Wasserstein convergence bounds, we show that optimal self-distillation improves the velocity estimation terms controlling continuous-time and finite-step generation error. Experiments with Gaussian models, Gaussian mixtures, and image data show that optimal self-distillation improves velocity risk, mode recovery, and finite-step generation relative to both the teacher and pure distillation.