Search papers, labs, and topics across Lattice.
This paper introduces AfriSwitch, a comprehensive benchmark for code-switched speech recognition in African languages, consisting of 61.36 hours of human-transcribed data. The study reveals significant variability in code-switching behavior across languages and demonstrates that existing ASR systems struggle with this complexity, achieving high word error rates (WER) that exceed those of monolingual benchmarks. Notably, performance is more closely linked to Africa-targeted training than to model size or nominal language coverage, highlighting the need for tailored approaches in multilingual ASR systems.
Existing ASR systems falter in recognizing code-switched speech, with the best performing system still achieving a staggering 35.93% word error rate.
Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), and switch-point counts. Corpus statistics show that mixing behaviour varies widely across African languages along two largely independent axes: how often speakers alternate, and how balanced the mixture is. No single scalar captures how code-switched a language is. Benchmarking five open and commercial multilingual ASR systems zero-shot yields word error rates far above published monolingual figures for the same languages, with the best system averaging 35.93% WER and no system falling below 24% on any language. Africa-targeted training, not model scale or nominal language coverage, best predicts performance.