Search papers, labs, and topics across Lattice.
Affiliation:
4
0
4
VIBE achieves superior instruction adherence and controllability in music generation from video, setting a new standard for semantic alignment in multimodal tasks.
Current VLMs miss crucial transition-level physics, but APT-Tune teaches them to learn causal transitions without forgetting event-level context.
Generate semantically aligned, high-fidelity music for videos with unprecedented speed and control by combining autoregressive planning and diffusion.
Video diffusion models can now generate physically plausible 4D worlds thanks to a new pipeline that combines pretraining, supervised fine-tuning, and reinforcement learning.