Search papers, labs, and topics across Lattice.
This paper introduces DeforM, a novel reasoning-guided framework for image-to-video generation that enhances the model's ability to synthesize physics-aware videos by focusing on critical dynamic regions. By integrating a VLM-guided physical reasoning module, DeforM-Reason, the framework effectively localizes and masks areas of interest, addressing the challenge of irrelevant region interference in video generation. Experimental results show that DeforM significantly improves both the visual quality and physical consistency of generated deformation scenarios compared to existing models.
DeforM boosts video generation realism by directing attention to physics-critical regions, outperforming traditional models in both quality and consistency.
Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be described by explicit trajectories or formulas, complex deformation dynamics remain challenging to synthesize. We observe that a lack of physical reasoning for localizing dynamic areas allows irrelevant regions to dilute the model's attention, leading to generation failure. In this paper, we propose DeforM, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions. To reason about and localize these critical regions, we introduce a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks. For physical guidance, we develop two alternative strategies: DeforM-Free for training-free mechanism analysis and DeforM-Injection as a powerful training-based generator. Experimental results demonstrate that DeforM improves the realism of generated deformation scenarios, outperforming baseline models in both visual quality and physical consistency.