Search papers, labs, and topics across Lattice.
This paper investigates On-Policy Delta Distillation (OPD$^2$) as a novel approach for enhancing multilingual mathematical reasoning in large language models (LLMs). By leveraging the probability gap between a post-trained teacher model and its base model, OPD$^2$ significantly outperforms traditional OPD, particularly in Korean and Japanese contexts, and reduces the performance disparity between English and Korean. The findings underscore the critical role of multilingual data in maintaining target-language performance while improving overall reasoning capabilities.
OPD$^2$ not only boosts multilingual math reasoning but also narrows the performance gap between English and Korean models, revealing the hidden potential of language-specific training signals.
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.