Search papers, labs, and topics across Lattice.
Affiliation:
6
5
4
Out-of-domain generalization in speech encoders hinges on their proximity to unseen TTS embeddings rather than their distance from natural speech.
Joint language-district supervision not only boosts district discrimination but also preserves language classification integrity, revealing a nuanced structure in speech embeddings.
FastConformer achieves over 90% accuracy on out-of-domain language identification, outperforming Whisper without task-specific adaptation.
ASR systems exhibit surprising language-specific sensitivities, revealing that speaker behavior and signal processing choices can drastically affect performance across Indic languages.
Geographic distance significantly predicts ASR performance, revealing that models struggle with regional variations in Indian languages.
VAANI's open-sourced dataset offers unprecedented coverage of India's linguistic landscape, finally giving researchers the data needed to build truly inclusive speech models.