Search papers, labs, and topics across Lattice.
7
0
7
13
Automated data curation and imbalance-aware training strategies significantly enhance LALMs' performance on culturally diverse folk music, yet deep musical understanding remains elusive.
Achieving a 76.2% relative gain in multi-step reasoning success rates, HDR redefines the capabilities of video models in real-time applications.
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Qwen-Music outperforms leading systems in musicality and audio quality, achieving state-of-the-art results across 13 of 16 metrics while generating songs from text and reinterpreting existing tracks.
Achieving high-fidelity audio generation with just four sampling steps, AudioX-Turbo dramatically cuts inference costs while enhancing performance across multimodal tasks.
TACO reduces token overhead by 10% while boosting terminal agent performance by up to 4%, revolutionizing how we approach long-horizon reasoning tasks.
Audio-Omni can edit sound, music, and speech with a single model, rivaling specialized systems and unlocking capabilities like knowledge-augmented reasoning and zero-shot cross-lingual control.