Search papers, labs, and topics across Lattice.
This paper introduces 4DSynth, a controllable procedural system that generates editable 4D environments from natural-language descriptions, blueprints, or photographs, integrating visual diversity and physical interactivity. The system allows for the creation of environments with explicit geometry and animated actors, while also enabling the generation of a scalable interactive navigation benchmark, 4DSynth-Nav. Evaluation of vision-language models on this benchmark reveals significant challenges, highlighting the importance of controllable environment generation for advancing embodied agent capabilities.
Vision-language models struggle significantly with interactive navigation tasks in procedurally generated environments, revealing critical gaps in current AI capabilities.
Embodied agents need environments that are visually diverse, physically interactive, and changing over time. Procedural simulators can generate large interactive scene collections, and recent 4D generators produce compelling visual dynamics. Combining these properties in one environment, however, still demands extensive manual effort, and the result is rarely editable or controllable enough to reuse at scale. We present 4DSynth, a controllable procedural system that turns a natural-language description, a blueprint mask, or a single photograph into an editable 4D environment with explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation state. Multiple scene routes share one geometry-grounded representation, so the same pipeline handles animation, camera planning, rendering, and task generation. To validate the full pipeline, we construct 4DSynth-Nav, an interactive navigation benchmark generated entirely from 4DSynth's procedural scenes. Two vision-language models evaluated across three difficulty tiers both fail the majority of tasks and stall after early subtasks. The same procedural controllability that produces these environments also makes each failure reproducible and each difficulty axis independently tunable. This paper presents both a controllable generation pipeline and the scalable benchmark it enables, offering a practical foundation for developing and evaluating embodied agents.