Search papers, labs, and topics across Lattice.
4
0
5
A unified taxonomy of audio editing tasks reveals the transformative potential of foundation models in reshaping how we interact with sound.
Removing gold answer strings from rewritten contexts can cause F1 scores to plummet by up to 64 points, underscoring their critical role in retrieval-augmented QA performance.
SwanVoice leaps ahead in zero-shot TTS by nailing expressive, multi-speaker dialogue with a single model, finally bridging the gap between monologue quality and conversational coherence.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.