Search papers, labs, and topics across Lattice.
This paper introduces GraphVid, a graph-conditioned image-to-video generation model that facilitates interactive control of multi-object interactions through structured interaction graphs, addressing the limitations of traditional trajectory-based control methods. By curating the GraphVid-Bench dataset, which includes large-scale interaction-centric video data with relational annotations, the authors enable the training of models that are both interaction-aware and efficient. GraphVid outperforms existing motion-control methods, achieving significant reductions in FID and FVD metrics while enhancing PSNR and SSIM, demonstrating the effectiveness of structured semantic interfaces in video generation.
GraphVid achieves superior video quality and controllability with significantly less training data, revolutionizing how we can interact with multi-object video generation.
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.