Search papers, labs, and topics across Lattice.
This paper investigates the behavioral and representational differences between World-Action Models (WAMs) and Vision-Language-Action (VLA) policies in robotic manipulation, focusing on whether WAMs offer meaningful improvements beyond mere task success. Using a model-agnostic diagnostic framework, the authors analyze both behavioral dynamics and internal representations across various policies, revealing that while WAMs enhance object-level behavior and target selectivity, these benefits are architecture-dependent and come with increased inference costs. The study finds that sequential WAMs exhibit the most predictive structure, suggesting that future WAM designs should prioritize maintaining actionable future representations for efficient manipulation tasks.
WAMs can boost object-level behavior and target selectivity, but their effectiveness hinges on architecture and incurs higher inference costs.
Vision-language-action (VLA) policies and World-Action Models (WAM) represent two increasingly important paradigms for robotic manipulation. However, it remains unclear whether future prediction in WAMs leads to behaviorally meaningful improvements beyond final task success. In this paper, we ask whether WAMs merely add future prediction, or whether they change robot behavior and internal representations in ways that are actionable for control. We introduce a model-agnostic diagnostic framework that compares WAMs and VLAs through two complementary lenses: behavioral rollout analysis and sparse-autoencoder-based feature analysis. The behavioral protocol measures action dynamics consistency, target-object progress, distractor disturbance, and runtime cost. The feature-space protocol characterizes internal representations as memorized, reactive, or predictive, revealing whether models encode future-oriented structure. Across LIBERO and RoboTwin2.0, we evaluate 7 policies spanning direct VLAs and joint, sequential, and auxiliary WAMs. Our results show that success alone hides key differences: WAMs often improve object-level behavior and target selectivity, but their gains depend on architecture and incur higher inference cost. Sequential WAMs show the clearest predictive structure, while auxiliary and joint WAMs respectively compress or entangle future information. These findings suggest future directions for WAMs design to preserve behaviorally actionable future representations for efficient manipulation.