Search papers, labs, and topics across Lattice.
WorldRover introduces a scalable synthetic video data engine designed to generate richly annotated explorations of artist-built environments, addressing the limitations of real-world capture for training interactive models. By leveraging an Unreal Engine pipeline, it produces detailed sequences that include RGB, metric depth, camera trajectories, and action signals, enabling comprehensive long-range exploration. The resulting dataset, WorldRover-10M, facilitates the development of models capable of constructing and revisiting coherent representations of complex environments, significantly enhancing the training resources available for AI systems focused on world exploration.
WorldRover turns long-horizon world exploration into a scalable data-generation problem, providing a treasure trove of richly annotated sequences for training AI models.
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.