Search papers, labs, and topics across Lattice.
3
0
5
4
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
DiaScriber achieves unprecedented accuracy in multi-speaker scenarios, overcoming the challenges of overlapping speech and rapid transitions.
Achieving high-quality audio reconstruction at unprecedented speeds, Qwen-Audio-VAE encodes 32 minutes of audio in just 541 ms.