Search papers, labs, and topics across Lattice.
This study addresses the challenge of developing robust speaker verification (SV) models for real classroom environments, focusing on both children and adults amidst classroom noise. By adapting the WavLM-TDNN model and employing a two-stage training strategy that combines self-supervised learning with fine-tuning on limited annotated data, the authors achieve significant improvements in Equal Error Rate (EER). The results show a 23.99% reduction in EER compared to the ECAPA-TDNN baseline, highlighting the effectiveness of their approach in enhancing SV performance in educational settings.
Achieving a 23.99% reduction in speaker verification errors in noisy classroom conditions could revolutionize AI applications in education.
Developing speaker verification (SV) models that are robust to classroom noise and effective across both children and adult speakers is critical for AI tools supporting educational environments. In this study, we use a real-world English-speaking classrooms dataset containing partial speaker identity annotations, with most recordings remaining unlabeled. We adapt the WavLM-TDNN model for classroom SV, achieving average relative reductions in Equal Error Rate (EER) of 23.99% and 6.32% compared to the ECAPA-TDNN baseline and the ECAPA-TDNN model trained on classroom data, respectively. Additionally, we investigate two training strategies for SV in classroom settings: self-supervised learning (SSL) and a two-stage approach that first pre-trains with SSL and then fine-tunes with limited annotated data. Five-fold cross-validation demonstrates that the two-stage strategy consistently outperforms the SSL-only approach, achieving an average relative EER reduction of 13.39%.