Search papers, labs, and topics across Lattice.
4
0
6
10
Compressing the KV cache of speech tokens can enhance decoding speed by over 1.49 times while improving performance on key benchmarks.
Interleaving speech and text during ASR training boosts entity recognition accuracy and narrows the gap between modalities, challenging traditional training paradigms.
Strong translation quality doesn't guarantee high speech or temporal fidelity, revealing critical gaps in existing evaluation practices for speech translation systems.
Unleashing LLMs' reasoning powers on speech unlocks a new ASR paradigm, slashing error rates by up to 17% simply by having the model "think" before transcribing.