Search papers, labs, and topics across Lattice.
This paper critiques traditional autoregressive video distillation methods that separate initialization and distribution matching stages, leading to suboptimal performance due to misalignment in target distributions. By introducing a distributional evaluation protocol, the authors reveal that high precision in initializations often comes at the cost of mode coverage, which is crucial for effective distillation. The proposed joint distillation method integrates mode-seeking and mode-covering objectives, resulting in improved generation quality, coverage, and diversity, even outperforming larger models in certain scenarios.
Achieving better video distillation quality isn't just about precision; it's about ensuring broad mode coverage during training.
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.