Search papers, labs, and topics across Lattice.
CG-World is a comprehensive dataset designed to capture the intricate dynamics of world models by integrating multimodal semantics, spatial structures, and various state parameters derived from industrial computer graphics. It comprises around 850,000 temporally aligned segments that facilitate the study of intervention learning and counterfactual reasoning through a structured approach to recording factual and alternative outcomes. Evaluation results indicate that CG-World enhances geometry-conditioned video generation, action prediction, and vision-language-action policy transfer, establishing a foundation for future research in Physical AI and embodied intelligence.
CG-World reveals that structured supervision from a vast array of world-state data can significantly improve action modeling and policy transfer in AI systems.
World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.