Search papers, labs, and topics across Lattice.
5
0
5
6
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
DiaScriber achieves unprecedented accuracy in multi-speaker scenarios, overcoming the challenges of overlapping speech and rapid transitions.
Achieving a 19.4% reduction in Mel Distance while using 45% fewer parameters, ear-VAE2 redefines the standards for high-fidelity music reconstruction.
Qwen-Music outperforms leading systems in musicality and audio quality, achieving state-of-the-art results across 13 of 16 metrics while generating songs from text and reinterpreting existing tracks.
Achieving high-quality audio reconstruction at unprecedented speeds, Qwen-Audio-VAE encodes 32 minutes of audio in just 541 ms.