Search papers, labs, and topics across Lattice.
This paper introduces CASA, a novel architecture that integrates Whisper-medium and Qwen3.5-2B for automatic speaking assessment (ASA), achieving state-of-the-art performance with a root mean square error (RMSE) of 0.358 on the Speak & Improve Corpus 2025. The approach allows for a clear distinction between speech delivery and content, enhancing interpretability while using approximately half the parameters of previous models. Through extensive ablation studies, the authors reveal the contributions of acoustic and content features, highlighting the model's robustness and adaptability across different ASA datasets.
CASA not only sets a new benchmark in automatic speaking assessment but also clarifies how acoustic and content features interact, paving the way for more interpretable AI assessments.
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.