Search papers, labs, and topics across Lattice.
This paper introduces an end-to-end pipeline for automated pronunciation evaluation of Korean toddler speech, addressing the significant gap in tools for assessing speech sound disorders in this demographic. By integrating neural speaker diarization with self-supervised learning, the authors present a novel corpus of recordings and evaluate various models, ultimately achieving high speaker count accuracy and low diarization error rates. The cross-model ensemble approach yields balanced accuracies for consonant and vowel predictions, demonstrating the effectiveness of combining different SSL backbones for improved performance in speech assessment.
Achieving 88.69% speaker count accuracy and a mean pronunciation scoring accuracy of 0.782 could revolutionize automated speech evaluation for Korean toddlers.
Speech sound disorders affect approximately 44% of Korean pediatric communication disorder cases, yet automated assessment tools for Korean toddler speech remain underdeveloped. This paper presents an end-to-end pipeline for automated pronunciation evaluation of Korean toddler speech, combining neural speaker diarization with self-supervised speech representation learning. We introduce a novel IRB-approved corpus of 53 recordings from Korean-speaking children aged 2-5 years. A subset of 53 subjects was annotated by three independent reviewers, yielding 1,190 consonant and 748 vowel word-level binary correctness labels. We evaluate three diarization models, finding that NeMo SortFormer achieves 88.69% speaker count accuracy and 33.04% diarization error rate (DER) owing to its arrival-time-sorted transformer architecture, which handles the acoustic confound between young female caregivers exhibiting aegyo and toddler speech. For pronunciation scoring, we compare three self-supervised learning (SSL) backbones across multiple pooling strategies. A cross-model ensemble routing consonant prediction to HuBERT-large and vowel prediction to WavLM-large achieves balanced accuracies of 0.720 and 0.845, with a mean of 0.782.