Search papers, labs, and topics across Lattice.
4
0
5
5
Audio-language models struggle significantly with hidden evaluation tasks, with accuracy plummeting by nearly 12 percentage points on average, highlighting the challenge of true audio comprehension.
Achieving one-step audio waveform generation with a 17脳 speedup while preserving quality could revolutionize TTS systems.
Speech QA performance peaks at 4.17 Hz, challenging the assumption that higher frame rates always yield better reasoning outcomes.
Unlock SOTA audio understanding by jointly training on readily available clip-level descriptions and scarce frame-level annotations, bridging the gap between global semantics and local details.