Search papers, labs, and topics across Lattice.
8
0
5
9
CSAVocoder achieves real-time spatial audio generation with enhanced fidelity by effectively integrating dynamic spatial cues, outperforming traditional vocoders.
VoxAudio revolutionizes vocalized audio synthesis by embedding intelligible speech seamlessly within complex soundscapes, outperforming traditional methods that compromise on clarity and control.
Achieving high-quality audio reconstruction at unprecedented speeds, Qwen-Audio-VAE encodes 32 minutes of audio in just 541 ms.
Qwen-Music outperforms leading systems in musicality and audio quality, achieving state-of-the-art results across 13 of 16 metrics while generating songs from text and reinterpreting existing tracks.
A unified taxonomy of audio editing tasks reveals the transformative potential of foundation models in reshaping how we interact with sound.
Spatial-Omni achieves superior spatial audio understanding by seamlessly integrating FOA encoding into existing LLMs, outperforming traditional models without compromising general audio processing.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.
Current speech generation models still fall short in maintaining consistency and capturing nuanced expressiveness when generating long-form speech, despite advances in high-fidelity synthesis.