Search papers, labs, and topics across Lattice.
Affiliation:
7
5
4
2
Out-of-domain generalization in speech encoders hinges on their proximity to unseen TTS embeddings rather than their distance from natural speech.
Aligning audio with image representations can drastically boost ASR performance in low-resource languages without the need for expensive transcriptions.
Joint language-district supervision not only boosts district discrimination but also preserves language classification integrity, revealing a nuanced structure in speech embeddings.
Incorporating synthetic speech data can lead to substantial performance improvements in ASR systems for Indic languages, but the choice of synthesis model and voice cloning strategy is critical.
Geographic distance significantly predicts ASR performance, revealing that models struggle with regional variations in Indian languages.
ASR systems exhibit surprising language-specific sensitivities, revealing that speaker behavior and signal processing choices can drastically affect performance across Indic languages.
VAANI's open-sourced dataset offers unprecedented coverage of India's linguistic landscape, finally giving researchers the data needed to build truly inclusive speech models.