Search papers, labs, and topics across Lattice.
To demystify the monolithic design of World-Action Models (WAMs), OpenWAM factorizes the architecture into modular components to systematically ablate how video-generative priors transfer to robotic control. The investigation reveals that effective cross-domain transfer requires dedicated action capacity, explicit world-to-action information flow, synchronized joint denoising, and single-stage co-training on egocentric human and robot data to drive out-of-domain generalization. Using these design principles, the resulting OpenWAM-$\alpha$ model is pretrained on roughly 6,400 hours of data and establishes top-tier control performance across eight simulation suites and diverse real-world manipulation platforms.
Video generative priors only translate into robust physical control when paired with explicit world-to-action information routing and synchronized joint denoising rather than standard monolithic fine-tuning.
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-{\alpha}, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-{\alpha} delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.