Search papers, labs, and topics across Lattice.
7
0
4
Correctly aligning audio descriptions with representations can boost classification performance by nearly 5 points, revealing the critical role of semantic refinement in audio tasks.
Integrating semantic understanding with acoustic fidelity can dramatically elevate the accuracy of speech quality assessments, challenging the dominance of self-supervised learning models.
FullDiT not only outperforms leading commercial music generators but also redefines how we approach music rendering by leveraging full-context generation techniques.
Unified generation of temporally structured audio achieves unprecedented speaker similarity and cross-turn consistency without task-specific branches.
Achieving high-quality full-length song generation from diverse inputs, this framework outperforms existing methods in musicality and fidelity.
Seamless transitions between speech and singing modes are now driven purely by text context, achieving state-of-the-art results in code-switching synthesis.
Forget clunky pipelines: this multi-agent system crafts compelling short dramas from a single sentence, nailing narrative pacing and spatial consistency in ways LLMs alone can't.