Search papers, labs, and topics across Lattice.
6
0
5
4
Achieving superior simultaneous speech translation quality without altering LLM architecture could redefine efficiency in real-time translation tasks.
Compressing the KV cache of speech tokens can enhance decoding speed by over 1.49 times while improving performance on key benchmarks.
Interleaving speech and text during ASR training boosts entity recognition accuracy and narrows the gap between modalities, challenging traditional training paradigms.
PRIME-Speech achieves low-latency, accurate speech-to-speech generation without sacrificing the robust performance of existing speech-to-text models.
Encoder-free speech modeling can rival traditional methods, challenging the necessity of dedicated speech encoders in LLM architectures.
Unleashing LLMs' reasoning powers on speech unlocks a new ASR paradigm, slashing error rates by up to 17% simply by having the model "think" before transcribing.