Search papers, labs, and topics across Lattice.
6
0
5
9
Qwen-Music outperforms leading systems in musicality and audio quality, achieving state-of-the-art results across 13 of 16 metrics while generating songs from text and reinterpreting existing tracks.
Achieving high-quality audio reconstruction at unprecedented speeds, Qwen-Audio-VAE encodes 32 minutes of audio in just 541 ms.
A unified taxonomy of audio editing tasks reveals the transformative potential of foundation models in reshaping how we interact with sound.
Spatial-Omni achieves superior spatial audio understanding by seamlessly integrating FOA encoding into existing LLMs, outperforming traditional models without compromising general audio processing.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.
Current speech generation models still fall short in maintaining consistency and capturing nuanced expressiveness when generating long-form speech, despite advances in high-fidelity synthesis.