Search papers, labs, and topics across Lattice.
This paper introduces VIBE, a text-and-video-to-music generation model that addresses the limitations of existing video-to-music models by incorporating a depth-wise cross-layer conditioning mechanism and a comprehensive reward modeling taxonomy. The model optimizes for both hard constraints like tempo and soft qualities such as musicality, employing a structured five-stage training curriculum. Evaluation results indicate that VIBE significantly improves controllability and adherence to instructions while maintaining competitive performance in generation fidelity and multimodal alignment compared to existing baselines.
VIBE achieves superior instruction adherence and controllability in music generation from video, setting a new standard for semantic alignment in multimodal tasks.
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Autoregressive (DAR) architectures. To resolve this, we introduce VIBE, a novel text-and-video-to-music (T+V2M) generation model that leverages: (1) Conditioning Connection, a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and (2) a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints (e.g., tempo, key) and soft, subjective qualities (e.g., musicality, multimodal alignment) with a structured 5-stage training curriculum. Upon evaluation using audio-visual alignment, instruction following, and audio quality metrics, along with a subjective human evaluation study, we observe that VIBE demonstrates enhanced controllability and instruction adherence while performing comparably to most evaluated baselines on generation fidelity and multimodal alignment.