Search papers, labs, and topics across Lattice.
3
0
4
2
Interleaving speech and text during ASR training boosts entity recognition accuracy and narrows the gap between modalities, challenging traditional training paradigms.
Encoder-free speech modeling can rival traditional methods, challenging the necessity of dedicated speech encoders in LLM architectures.
Unleashing LLMs' reasoning powers on speech unlocks a new ASR paradigm, slashing error rates by up to 17% simply by having the model "think" before transcribing.