Search papers, labs, and topics across Lattice.
This paper introduces Multi-Teacher Self-Distillation Policy Optimization (MT-SDPO), an innovative approach that enhances the integration of multiple domain-specific teachers into a single large language model (LLM) by dynamically selecting the most reliable teacher for each sample. The method leverages self-anchors, answer verification, and privileged distillation to ensure that supervision is based on verified correctness rather than fixed domain labels. Results show that MT-SDPO significantly improves the performance of the Qwen3-8B model, increasing its weakest domain score by 14.79 points and reducing the domain gap by 74.7%, demonstrating the efficacy of verified reliability in teacher selection.
Verified reliability outperforms domain expertise in teacher selection, leading to substantial performance gains in multi-domain LLMs.
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose domain matches it, existing approaches let a domain label decide which teacher provides supervision. However, domain expertise holds only on average: the matched teacher is not always correct on a given sample, while a teacher from another domain sometimes is. The reliable teacher therefore has to be identified per sample, not per domain. In this paper, we introduce Multi-Teacher Self-Distillation Policy Optimization (MT-SDPO), an on-policy distillation method that unifies several frozen teachers into one student model. MT-SDPO consists of three components: (1) self-anchors, where a rollout is supervised by a correct rollout from its own group; (2) answer-verified eligibility, where a teacher may supervise a sample only if its own answer passes a verifier; and (3) privileged distillation, which merges the anchor and all verified feedback into one context that an exponential moving average self-teacher reads and the student does not, thereby keeping one policy at deployment. Across five students from three model families, MT-SDPO lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain. Verified reliability, not domain membership, should decide who teaches. Code is available at https://github.com/hexixiang/MT-SDPO.