Search papers, labs, and topics across Lattice.
This paper introduces DECOWAM, a decoupled whole-body world-action model designed for legged mobile manipulation, which effectively separates camera ego-motion from base and arm actions using dedicated conditional interfaces. By utilizing an adapted FastWAM backbone and training residual adapters, DECOWAM achieves a 21.7% reduction in action mean squared error (MSE) while maintaining task completion rates comparable to strong baselines. The introduction of the ARMDOG dataset further enhances the model's performance by providing synchronized video and state-action data, demonstrating improved whole-body coordination and robustness in dynamic environments.
DECOWAM achieves a 21.7% reduction in action prediction error while maintaining robust task performance, showcasing the power of embodiment-aware factorization in mobile manipulation.
Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.