Search papers, labs, and topics across Lattice.
This study examines how the linguistic distance of speakers' first languages (L1) from English affects their performance in automatic speech recognition (ASR) systems. By employing Tweedie mixed-effects models, the authors establish a statistically significant correlation between L1 distance and ASR error rates, revealing that this effect varies across different datasets and models. Furthermore, the analysis of latent representations indicates that deeper acoustic layers exhibit L1-based spatial segregation, highlighting inherent biases in ASR systems that could impact diverse speaker populations.
ASR systems show systematic performance disparities linked to the linguistic distance of speakers' first languages, revealing hidden biases in their design.
While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers' L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models ($p<0.001$ across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures