Search papers, labs, and topics across Lattice.
11
0
7
8
VoxAudio revolutionizes vocalized audio synthesis by embedding intelligible speech seamlessly within complex soundscapes, outperforming traditional methods that compromise on clarity and control.
SwanTale achieves unprecedented expressiveness in multi-speaker audio generation, outperforming existing models in both instruct and zero-shot tasks.
A unified taxonomy of audio editing tasks reveals the transformative potential of foundation models in reshaping how we interact with sound.
Full-duplex dialogue systems are often mischaracterized, with many claiming capabilities they cannot deliver due to training limitations.
Spatial-Omni achieves superior spatial audio understanding by seamlessly integrating FOA encoding into existing LLMs, outperforming traditional models without compromising general audio processing.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.
SwanVoice leaps ahead in zero-shot TTS by nailing expressive, multi-speaker dialogue with a single model, finally bridging the gap between monologue quality and conversational coherence.
Current speech generation models still fall short in maintaining consistency and capturing nuanced expressiveness when generating long-form speech, despite advances in high-fidelity synthesis.
Current audio-visual models nail unimodal quality but still struggle to make music and dance move together rhythmically, highlighting a key gap TMD-Bench is designed to address.
Turns out, your image-generating diffusion model already knows how to segment anything you ask it to.
Current reward models for spoken dialogue systems are missing crucial paralinguistic and natural speech elements, but this new model closes the gap by operating directly on speech and outperforming existing audio LLMs.