Search papers, labs, and topics across Lattice.
This paper introduces VideoCoCo, an innovative dual-engine framework that utilizes executable Blender code as a process-level chain of thought to enhance physically consistent video generation from text prompts. By synthesizing a Blender program that specifies scene dynamics and employing a generative video engine for photorealistic rendering, VideoCoCo effectively separates process-level reasoning from visual realization. The approach significantly outperforms existing methods, improving scores on PhyGenBench and VBench-2.0, thus showcasing the efficacy of executable code in achieving controllable and inspectable video generation.
Executable Blender code transforms text-to-video generation, enabling unprecedented control over scene dynamics and visual fidelity.
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.