Search papers, labs, and topics across Lattice.
4
0
5
8
Audio-language models struggle significantly with hidden evaluation tasks, with accuracy plummeting by nearly 12 percentage points on average, highlighting the challenge of true audio comprehension.
Current audio editing models are failing spectacularly, with an Exact Match Rate below 5% in complex tasks, exposing a critical need for improvement.
Unlock SOTA audio understanding by jointly training on readily available clip-level descriptions and scarce frame-level annotations, bridging the gap between global semantics and local details.
Online reinforcement learning with large audio language model rewards catapults text-to-audio generation to a new state-of-the-art, even with a relatively small 470M parameter model.