Search papers, labs, and topics across Lattice.
This study evaluates the recognition performance of human listeners versus three advanced automatic speech recognition (ASR) systems on Dutch dysarthric speech, revealing that both groups struggle with word error rates (WER) exceeding 70%. Notably, fine-tuning ASR models on dysarthric speech significantly reduced WER, with personalized models outperforming human listeners despite still high overall error rates. These findings highlight the potential for tailored DSR models to enhance communication for individuals with severe dysarthria, moving closer to practical applications in everyday settings.
Personalized dysarthric speech recognition models can outperform human listeners, despite both facing significant challenges in understanding severe dysarthric speech.
In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and those of three state-of-the-art, off-the-shelf ASR systems (Whisper-large-V3, Google Chirp 3, and Omnilingual) on the recognition of Dutch continuous read and spontaneous speech from a single speaker with severe dysarthria. Results showed that both humans listeners and the three off-the-shelf ASR systems exhibit word error rates (WER) exceeding 70% on average, indicating that DSR is highly challenging for both humans and ASR systems. Fine-tuning on the dysarthric speech significantly reduced WER. Although overall WERs are still quite high (>23%), the personalised DSR models outperformed the human listeners, and performance is getting closer to being useful for supporting day-to-day communication of dysarthric speakers. Future research should focus on improving personalized DSR on spontaneous speech and longer utterances in the case of read speech, with a specific focus on particular phonemes.