Search papers, labs, and topics across Lattice.
This study establishes a reproducible multi-seed ASR benchmark for the low-resource Garhwali language, revealing that previously reported gains in ASR performance are often fragile and not replicable. The authors found that neither advanced objectives like Focal CTC nor matra-weighted approaches outperformed standard CTC when subjected to rigorous seed-level testing, and that Hindi-to-Garhwali transfer learning provided no significant advantage over direct fine-tuning. Ultimately, the best performance was achieved with w2v-BERT 2.0 using standard CTC, highlighting that pretraining design is more critical than model size in this context, with a reported 47.0% WER across five seeds.
Fragile gains in low-resource ASR performance are exposed, with standard methods outperforming complex alternatives in Garhwali speech recognition.
At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the official VAANI splits, with per-seed outputs and significance testing. Re-examining plausible gains, we find them fragile: neither Focal CTC nor a matra-weighted objective beats standard CTC under seed-level testing, the matra objective fails to cut even its targeted errors, and Hindi-to-Garhwali transfer gives no gain over direct fine-tuning. What holds up is mundane: w2v-BERT 2.0 with standard CTC reaches 47.0% WER over five seeds, beating the larger MMS-1B and comparable models; pretraining design, not parameter count, drives performance, and speed augmentation gives a small, largely consistent gain. Multi-seed evaluation on official splits separates real gains from seed noise.