EmbodyX IncNortheasternUniversity of GeorgiaJun 3, 2026arXiv:2606.05254

Flash-WAM: Modality-Aware Distillation for World Action Models

Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang

AI Summary

This paper introduces Flash-WAM, a modality-aware step-distillation framework that optimizes the training of world-action models (WAMs) by addressing the asymmetry in noise distributions between video and action streams. By employing a tailored consistency function for each modality, Flash-WAM achieves a remarkable 23-fold reduction in inference latency, compressing it from 8.1 seconds to just 348 milliseconds while maintaining high task success rates. The framework not only preserves performance on simulation benchmarks but also significantly enhances real-world action execution on a humanoid robot, demonstrating its practical viability for real-time control applications.

Key Contribution

Flash-WAM achieves real-time inference for world-action models by reducing latency from 8.1 seconds to 348 milliseconds without sacrificing performance.

Abstract

World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce Flash-WAM, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition. Instantiated on LingBot-VA, Flash-WAM compresses inference to a single step in each modality. On RoboTwin 2.0, this reduces per-chunk latency from 8.1 seconds to 348 ms on NVIDIA L40S, a 23{times} speedup that enables real-time inference. Flash-WAM preserves task success on simulation benchmarks (85.5% RoboTwin 2.0, 95.7% LIBERO) and substantially recovers real-world performance (60% average on a Unitree G1 humanoid robot), while naive consistency distillation drops to 24% at the same step budget.

Multimodal Models Robotics & Embodied AI World Models & Planning

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Flash-WAM: Modality-Aware Distillation for World Action Models

Related Papers