Search papers, labs, and topics across Lattice.
This study explores the use of Dynamic Time Warping (DTW) on self-supervised WavLM representations to assess phonetic accuracy, rhythm, and intonation in L2 speech without relying on labeled data. The findings reveal that this DTW-based approach not only surpasses human agreement in holistic and sentence-level phonetic scoring but also achieves near-human performance in rhythm assessment. While the method shows promise for multi-aspect pronunciation evaluation, intonation scoring remains less robust, indicating areas for further improvement.
A DTW-based framework using self-supervised representations outperforms human raters in assessing L2 phonetic accuracy and rhythm, challenging traditional assessment methods.
L2 speech assessment has traditionally focused on phonetic assessment, leaving the scoring of suprasegmental features such as rhythm and intonation underexplored. Moreover, assessment methods often require training with labeled L2 speech data, making them difficult to apply in low-resource settings. We investigate whether DTW over self-supervised WavLM representations can provide a text-free framework for assessing phonetic accuracy, rhythm, and intonation in English and Japanese L2 speech. Results show that a basic DTW-based approach that compares learner speech to native templates exceeds human agreement on holistic and sentence-level phonetic scoring. For rhythm, we introduce methods that measure the degree of warping in the DTW alignment path; our best method approaches human-level performance. For intonation, we combine DTW distance over prosodic residuals with pitch and intensity features, but performance remains more modest on some tasks. Our results point to self-supervised representations as a promising, text-free basis for multi-aspect pronunciation assessment.