Search papers, labs, and topics across Lattice.
This paper critiques the technical failures of automatic speech recognition (ASR) systems in representing low-resource, Indigenous, and non-standard language varieties, framing these issues as reflections of colonial language hierarchies rather than mere technical shortcomings. By introducing the Three Harms (3M) taxonomy and a seven-layer situatedness model, the authors highlight how existing data, metrics, and model priors systematically marginalize certain voices. The proposed participatory framework aims to empower affected communities in the design and evaluation of ASR technologies, fostering a more inclusive approach to linguistic diversity in voice interfaces.
ASR systems are not just failing technically; they perpetuate colonial hierarchies that silence marginalized voices, necessitating a radical rethinking of how we design these technologies.
This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.