Search papers, labs, and topics across Lattice.
This study introduces Whisper-based systems for speaker-role and language diarization to analyze multilingual interviews aimed at assessing the language proficiency of older adults. By leveraging these systems, the authors demonstrate that language-adapted models significantly enhance diarization accuracy for lower-resource Indian languages, while statistical analyses identify key conversational behaviors as strong indicators of proficiency. The findings reveal that simple diarization-derived features can match the performance of advanced speech embeddings for proficiency prediction, highlighting the effectiveness of automated conversational analysis in language assessment.
Diarization-derived features can rival complex speech embeddings in predicting language proficiency, making automated assessments more accessible and efficient.
Automatic language proficiency assessment in the context of multilingual interview-based settings remains underexplored. In this work, we develop Whisper-based speaker-role and language diarization systems to automatically extract respondent speech and characterize language usage in multilingual interviews with older adults. We further investigate whether diarization-derived conversational and language-use behaviors can support downstream language proficiency assessment. Results show that language-adapted Whisper models substantially improve language diarization performance for lower-resource and linguistically related Indian languages. Statistical analyses reveal that respondent speech ratio and intended language usage are strong predictors of proficiency ratings. Furthermore, simple diarization-derived behavioral features achieve performance comparable to Whisper-based speech embeddings for proficiency prediction, while combining both yields the best results. Importantly, both the speech and language use statistical analyses and language proficiency prediction performance remain largely preserved when using fully automatic diarization outputs, demonstrating the potential of respondent-centric conversational analysis for scalable language proficiency assessment.