Search papers, labs, and topics across Lattice.
This paper investigates the robustness of World Action Models (WAMs) by leveraging mechanistic interpretability to analyze how robustness-relevant perturbations manifest in WAM activation space. The authors identify that certain WAM architectures demonstrate low-dimensional linear separability for critical features, which informs the development of a training-free steering method using contrastive activation directions. The proposed World-Action Linear Quadratic Regulator (WA-LQR) effectively enhances the robustness of Cosmos-Policy and DiT4DiT models against various perturbations, outperforming both unsteered and prompt steering baselines.
Some World Action Models can be steered to enhance robustness against perturbations without additional training, revealing a new avenue for improving model reliability.
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.