Search papers, labs, and topics across Lattice.
This study developed an Automatic Speech Recognition (ASR) system for Mizo, a low-resource language, by collecting 17.62 hours of speech data and fine-tuning both Whisper multilingual models and the SraVaani 1.0 Indic multilingual model. The Whisper-large-v3 model achieved the lowest conventional word error rate (WER) of 18.08%, while a morphology-aware evaluation further improved performance to a WER of 7.22%. In contrast, the SraVaani model showed a significant performance boost from zero-shot evaluation to fine-tuning, reducing its conventional WER from 58.27% to 29.45% and its morphology-aware WER to 17.93%.
Whisper's fine-tuning can reduce Mizo ASR error rates to as low as 7.22%, showcasing its potential for low-resource languages.
This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.