Search papers, labs, and topics across Lattice.
4
0
6
7
Transforming audio captioning from a passive task into an adaptive, evidence-driven process could redefine how we approach fine-grained audio understanding.
Fine-grained audio captioning just got a major upgrade鈥擜udioMap achieves state-of-the-art results by redefining how we reward temporal accuracy and descriptive richness in audio events.
Fine-grained evaluation reveals that current text-to-audio models fail to preserve speech content and control audio attributes effectively.
Forget supervised fine-tuning: RL alone can unlock high-quality chain-of-thought reasoning in audio-language models, even starting from a model with no prior CoT capability.