Search papers, labs, and topics across Lattice.
This paper addresses the issue of speaker leakage in cascaded multi-talker automatic speech recognition (MT-ASR) systems, which limits their performance despite the use of advanced foundation models. The authors introduce a pruning-based approach that leverages a pre-trained speaker diarization model to effectively identify and eliminate leakage artifacts by ensuring a tripartite consensus among temporal containment, lexical cross-validation, and temporal alignment. Experimental results demonstrate significant improvements, with up to 29% relative reductions in corrected word error rate (cpW ER) in scenarios with high speaker leakage, underscoring the method's robustness in challenging acoustic environments.
Pruning-based correction can cut speaker leakage in multi-talker ASR by up to 29%, transforming transcript reliability in noisy settings.
While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during separation. Prior correction strategies primarily focus on lexical re-labeling for speaker attribution. We propose a complementary pruning-based paradigm that robustly identifies and removes leakage artifacts. Our method utilizes a pre-trained speaker diarization model as a multimodal verifier to prune transcribed segments satisfying a tripartite consensus of temporal containment, lexical cross-validation, and temporal alignment. Results on LibriMix, LibriSpeechMix, and the AMI Meeting corpus show our algorithm consistently reduces cpW ER across diverse overlap conditions. Specifically, on subsets with high speaker leakage, our method achieves relative cpW ER reductions of up to 29%, highlighting its effectiveness in enhancing the reliability of cascaded MT-ASR transcripts in complex acoustic environments.