Search papers, labs, and topics across Lattice.
This paper introduces SNM-VFI, a training-free framework for motion-controllable generative video frame interpolation that leverages pre-trained optical flow and video diffusion models. By utilizing a symmetric nonlinear motion model, the method constructs multi-frame flow-based intermediate frames that guide the generative process, enhancing both perceptual realism and motion correspondence. Extensive evaluations on benchmarks such as DAVIS, Sintel, and KITTI reveal that SNM-VFI outperforms existing methods in terms of perceptual quality, reconstruction accuracy, and temporal coherence across various motion scenarios.
SNM-VFI achieves superior perceptual quality and temporal coherence in video frame interpolation by effectively integrating flow-guided frames with diffusion-generated details.
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model. Specifically, we first utilize a pre-trained optical flow model to construct multi-frame nonlinear flow-based intermediate frames and confidence maps. These flow-guided frames are then encoded as latent priors to initialize and iteratively guide a pre-trained Video Diffusion model, enabling the diffusion model to preserve dense motion correspondence while improving perceptual realism. To further enhance output quality, we employ confidence maps to fuse structurally reliable flow-based predictions with diffusion-generated details in uncertain regions such as occlusions and object boundaries. Extensive evaluations on challenging benchmarks, including DAVIS, Sintel, and KITTI, demonstrate that SNM-VFI achieves strong perceptual quality, competitive reconstruction accuracy, and robust temporal coherence across diverse motion scenarios.