Search papers, labs, and topics across Lattice.
This paper critiques the conventional approach of scaling world models through increased video data and compute, arguing that a more effective strategy involves a recursive data engine that provides grounded reward signals. It highlights the advantages of game development as a robust environment for generating high-quality rewards for Reinforcement Learning (RL), contrasting it with the limitations of current fuzzy proxies like CLIP scores. The proposed Reinforcement Learning with Human-Engine Verification (RLHEV) paradigm leverages executable game environments to deliver dense feedback and long-horizon trajectory data, significantly enhancing the post-training phase for spatial world models.
Game development offers a powerful alternative to traditional reward signals, enabling RL systems to leverage executable environments for superior feedback and data generation.
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.