Search papers, labs, and topics across Lattice.
8
0
8
18
DiaScriber achieves unprecedented accuracy in multi-speaker scenarios, overcoming the challenges of overlapping speech and rapid transitions.
Achieving over 90% performance retention with a staggering 20x KV cache compression could redefine efficiency in long-context audio inference.
Bridging the reasoning gap, X$^3$-OPD enables audio-language models to outperform their text-based counterparts in logical reasoning tasks.
Qwen-Music outperforms leading systems in musicality and audio quality, achieving state-of-the-art results across 13 of 16 metrics while generating songs from text and reinterpreting existing tracks.
Automatically constructed data can dramatically enhance the temporal localization abilities of audio models, overcoming the limitations of manual annotation.
Full-duplex dialogue systems are often mischaracterized, with many claiming capabilities they cannot deliver due to training limitations.
Audio LLMs can now be systematically evaluated for character alignment in role-playing scenarios, thanks to a new framework that judges both text and vocal features.
Current reward models for spoken dialogue systems are missing crucial paralinguistic and natural speech elements, but this new model closes the gap by operating directly on speech and outperforming existing audio LLMs.