Search papers, labs, and topics across Lattice.
This study investigates the effects of logit-based knowledge distillation (KD) during mid-training of language models, revealing that while forward KD enhances reasoning and factual recall in pre-training, it hampers factual recall during mid-training despite continued reasoning improvements. The authors attribute this phenomenon to the asymmetry in teacher confidence across different data types and the evolving knowledge state of the student model. To address this issue, they introduce Switch Distillation, which selectively distills knowledge based on teacher confidence, resulting in significant performance gains in reasoning and knowledge retention while preserving factual recall.
Switch Distillation not only enhances reasoning performance by up to 71% but also preserves factual recall, challenging the conventional trade-off in knowledge distillation.
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.