Search papers, labs, and topics across Lattice.
This study introduces a tone-conditioned curriculum learning framework tailored for low-resource Southern Bantu languages, addressing the high word error rates (WER) that hinder automatic speech recognition (ASR) applications in these languages. By integrating hybrid difficulty scoring, gated adapters based on tonal statistics, and staged curriculum training, the authors trained models on a community corpus and evaluated their performance on the NCHLT dataset. The results indicate significant architecture-language interactions, with W2V-BERT achieving a 28.41% average WER, outperforming Whisper on Nguni languages while the latter excelled in Sotho-Tswana languages, highlighting the necessity of model selection based on language specifics.
W2V-BERT with tone conditioning slashed average word error rates to 28.41% for low-resource Southern Bantu languages, revealing critical architecture-language dynamics.
Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.