Search papers, labs, and topics across Lattice.
1Zhejiang University
5
0
8
Achieving over 90% performance retention with a staggering 20x KV cache compression could redefine efficiency in long-context audio inference.
Bridging the reasoning gap, X$^3$-OPD enables audio-language models to outperform their text-based counterparts in logical reasoning tasks.
Automatically constructed data can dramatically enhance the temporal localization abilities of audio models, overcoming the limitations of manual annotation.
Achieving sub-centimeter accuracy in trajectory-controlled text-to-motion generation without sacrificing the quality of the pretrained model could revolutionize animation workflows.
Audio LLMs can now be systematically evaluated for character alignment in role-playing scenarios, thanks to a new framework that judges both text and vocal features.