Search papers, labs, and topics across Lattice.
This paper introduces Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), a novel approach that separates language-specific knowledge acquisition from the integration of multilingual capabilities in automatic speech recognition (ASR) systems. By employing reinforcement learning to optimize language-specialized teachers and subsequently distilling their expertise into a generalist multilingual student, the method effectively mitigates optimization conflicts arising from heterogeneous language characteristics. Experimental results on Mandarin, Cantonese, and English benchmarks show that LS-MOPD significantly outperforms existing reinforcement learning baselines, indicating its superior ability to generalize across languages in ASR tasks.
Language-Specialized Multi-Teacher On-Policy Distillation outperforms traditional RL methods, revealing a new pathway for enhancing multilingual ASR performance.
Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs'cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, joint modeling of languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), after which their expertise is integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore two acoustic-prefix configurations, static and dynamic, to examine how teacher--student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and consistently surpasses the empirical performance envelope defined by best-performing RL teachers, revealing its potential to generalize beyond all teachers in multilingual ASR.