Search papers, labs, and topics across Lattice.
This paper introduces a novel method for Mispronunciation Detection and Diagnosis (MDD) that utilizes language-specific statistical graphs to model phoneme confusion patterns as directed graphs. By incorporating a strategy tailored to different native language backgrounds, the authors effectively capture systematic pronunciation variations. The proposed approach achieves a notable F1-score of 59.52% on the L2-ARCTIC benchmark, surpassing multiple competitive baselines, highlighting its potential for enhancing language learning technologies.
Language-specific statistical graphs reveal phoneme confusion patterns, achieving an F1-score of 59.52% in mispronunciation detection鈥攐utperforming existing methods.
Mispronunciation Detection and Diagnosis (MDD) has gained increasing importance in computer-assisted language learning and speech technology in recent years. In this paper, we propose a method for constructing statistical graphs that enable models to learn phoneme confusion patterns represented as directed graphs. Furthermore, we introduce a language-specific strategy to capture systematic pronunciation differences across various native language (L1) backgrounds. The effectiveness of our approach is demonstrated through extensive experiments on the L2-ARCTIC benchmark, where it achieves an F1-score of 59.52%, outperforming several competitive baselines.