Search papers, labs, and topics across Lattice.
This study benchmarks 11 automatic speech recognition (ASR) models using dual-reference evaluations鈥攙erbatim and intended transcriptions鈥攕pecifically focusing on atypical stuttered speech. The findings reveal significant discrepancies in model performance and rankings based on the chosen reference style, emphasizing that conflating these two types can mislead evaluations. By highlighting the impact of transcription choice on ASR accuracy, this research advocates for a more nuanced approach to model assessment in atypical speech contexts.
Choosing the right transcription reference can drastically change the perceived performance of ASR systems, especially in atypical speech scenarios.
ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (actual produced speech, including repetitions/prolongations) and intended (the canonical form of the text with disfluencies removed) in atypical speech recognition depending on context and use-case. Most ASR evaluations conflate this duality into a single ground truth and reward systems that delete disfluencies, ignoring verbatim faithfulness. We benchmark 11 ASR models from encoder-decoder, CTC and transducer families using both verbatim and intended references on atypical stuttered speech as a case study. Our quantitative assessment underlines the disparity in model performance and rankings using the two transcript styles. Through this analysis, we highlight the importance of selecting a suitable transcription reference for valid model selection depending on the use-case, particularly for atypical ASR.