Search papers, labs, and topics across Lattice.
This paper introduces DonorRank, a learning-to-rank framework designed to enhance donor language selection for zero-shot automatic speech recognition (ASR) in low-resource settings. By evaluating its performance on multilingual speech corpora from Indic and African language families, the authors demonstrate that DonorRank significantly outperforms traditional heuristics based on genetic similarity or resource availability. The findings reveal that the composition of the donor set critically influences the effectiveness of linguistic cue transfer, providing actionable insights for improving multilingual ASR systems.
DonorRank reveals that the right selection of donor languages can drastically enhance zero-shot ASR performance in low-resource contexts, challenging existing heuristics.
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.