Search papers, labs, and topics across Lattice.
7
0
5
5
SR-FD reduces word error rates by over a third, transforming the landscape of intelligibility in few-step TTS synthesis.
ORCA not only boosts performance by 26.4 points but also restores critical speaker identity cues that traditional models overlook.
Taiwanese Mandarin TTS systems can achieve a 63.9% reduction in word error rates by using a context-adapted tokenizer and language model.
Leveraging lecture context can boost technical term recognition in Mandarin ASR systems by over 15% without sacrificing overall accuracy.
LatentASR transforms frozen ASR models into adaptive systems that intelligently allocate compute resources, achieving significant WER reductions on challenging inputs without the need for extensive retraining.
Bridging the gap between audio reconstruction and language modeling objectives yields neural audio codecs that are both more acoustically faithful and linguistically predictable.
Overcome LALM's struggles with localized dialectal prosody: a new Taiwanese audio-text dataset and fine-tuning strategy boosts accuracy by 6.5% on the TAU Benchmark.