Search papers, labs, and topics across Lattice.
HarmoniDPO introduces a novel framework for video-to-audio (V2A) generation that addresses the challenges of temporal synchronization and perceptual quality by integrating preference-based optimization into diffusion models. By utilizing a dual video representation that combines global context with frame-wise features, the method preserves temporal dynamics and semantic detail, while employing online Direct Preference Optimization to align audio output with human preferences. Experimental results show that HarmoniDPO significantly outperforms existing methods in both audio-video synchronization and subjective audio quality, marking a substantial advancement in realistic audio generation from video inputs.
HarmoniDPO achieves superior audio-video synchronization and quality by leveraging human preferences in a novel diffusion-based framework.
Video-to-audio (V2A) generation faces significant challenges in achieving precise temporal synchronization and high perceptual quality due to the complex and ambiguous relationship between visual and auditory cues. Existing methods typically compress video inputs into single feature representations, leading to a significant loss of temporal dynamics and fine-grained visual information. These approaches also rely on reconstruction-based training objectives that poorly correlate with human perceptual judgments of audio quality and appropriateness. To address these limitations, We propose HarmoniDPO, a novel framework that integrates preference-based optimization into diffusion-based V2A generation. Specifically, (1) Our approach leverages a dual video representation-combining global context with frame-wise features-to preserve temporal dynamics and semantic detail. (2) Inspired by reinforcement learning from human feedback (RLHF), HarmoniDPO employs online Direct Preference Optimization (online DPO) to fine-tune a diffusion-based V2A model using preference judgments, thus enhancing both perceptual quality and alignment. (3) Additionally, we introduce Dual-scale Diffusion Search (DDS), a test-time scaling algorithm that adaptively optimizes output fidelity during inference. Experiments demonstrate that HarmoniDPO outperforms state-of-the-art methods in audio-video synchronization and subjective audio quality, offering a robust solution for generating realistic, human-preferred audio from video.