Search papers, labs, and topics across Lattice.
This study benchmarks the performance of state-of-the-art automatic speech recognition (ASR) systems against Dutch native listeners in recognizing diverse speech, including that of children and older adults. The findings reveal that Google Telephony outperformed other ASR systems and, in certain instances, surpassed human listeners, highlighting the competitive capabilities of ASR in understanding varied speech patterns. Notably, performance discrepancies were observed based on speaker age, regional accents, and utterance length, suggesting avenues for enhancing ASR robustness in these areas.
Google Telephony not only rivals but sometimes outperforms human listeners in recognizing diverse speech, challenging traditional views on ASR limitations.
Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and Dutch native listeners on the recognition of"diverse"speech, specifically Dutch child and older adults'speech and Flemish. Google Telephony outperformed the other ASR systems. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them. Slight performance differences between the listeners and ASR systems were found related to speaker's age and regional accents and utterance length. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents. A comparison of ASR recognition performances on the test stimuli and the full Jasmin-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance.