Search papers, labs, and topics across Lattice.
This paper reviews the evolution of World Action Models (WAMs) and identifies critical gaps in their application to open-world physical intelligence, including issues with model roles, objectives, and system composition. The authors propose a co-evolution roadmap centered on the concept of the "embodied brain," which aims to integrate multimodal context and facilitate more effective decision-making in physical environments. Key findings suggest that by leveraging WAMs as prototypes and establishing shared contracts across models and tasks, a modular stack for adaptive embodied agents can be developed, enhancing their ability to learn and improve from interactions.
Bridging the gaps in physical intelligence could enable agents to learn and adapt in real-world scenarios more effectively than ever before.
Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems expose limited interfaces for reuse and evaluation. We review the evolution toward WAMs and organize these limitations into three coupled gaps: model roles and representations, objectives and standardization, and system composition. Building on this analysis, we propose a co-evolution roadmap for physical intelligence centered on the \emph{embodied brain}, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands. WAMs provide promising prototypes for its predictive functions, while a physical harness grounds model outputs through tools, controllers, verification, and trace logging. Shared contracts align heterogeneous models, data, tasks, and embodiments, and closed-loop post-training converts verified interaction into reusable experience. Together, these components define a modular physical-intelligence stack for adaptive and self-improving embodied agents.