Search papers, labs, and topics across Lattice.
6
0
6
14
SwanTale achieves unprecedented expressiveness in multi-speaker audio generation, outperforming existing models in both instruct and zero-shot tasks.
A unified taxonomy of audio editing tasks reveals the transformative potential of foundation models in reshaping how we interact with sound.
Removing gold answer strings from rewritten contexts can cause F1 scores to plummet by up to 64 points, underscoring their critical role in retrieval-augmented QA performance.
SwanVoice leaps ahead in zero-shot TTS by nailing expressive, multi-speaker dialogue with a single model, finally bridging the gap between monologue quality and conversational coherence.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.
Current speech generation models still fall short in maintaining consistency and capturing nuanced expressiveness when generating long-form speech, despite advances in high-fidelity synthesis.