Search papers, labs, and topics across Lattice.
3
0
4
4
Audio-language models struggle significantly with hidden evaluation tasks, with accuracy plummeting by nearly 12 percentage points on average, highlighting the challenge of true audio comprehension.
Pruning can slash the computational cost of text-to-audio models by over 80% without sacrificing quality, but it poses risks to generating critical sound events.
Adapting speech enhancement models can significantly boost singing voice separation performance without sacrificing their original capabilities.