Search papers, labs, and topics across Lattice.
5
0
3
Correctly aligning audio descriptions with representations can boost classification performance by nearly 5 points, revealing the critical role of semantic refinement in audio tasks.
LoRA-GA$^2$ closes the performance gap with full fine-tuning by leveraging multi-step gradient dynamics without sacrificing efficiency.
FullDiT not only outperforms leading commercial music generators but also redefines how we approach music rendering by leveraging full-context generation techniques.
Unified generation of temporally structured audio achieves unprecedented speaker similarity and cross-turn consistency without task-specific branches.
Achieving high-quality full-length song generation from diverse inputs, this framework outperforms existing methods in musicality and fidelity.