Mar 8, 2026arXiv:2603.07648

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

Likui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen, Zisheng Chen, Jianhua Han, Jiangtong Zhu, Pei Xu, Hang Xu, Hefeng Wu, Liang Lin, Xiaodan Liang

AI Summary

AtomicVLA is introduced as a novel framework for robotic manipulation that decomposes complex tasks into task-level plans, atomic skill abstractions, and fine-grained actions. It employs a Skill-Guided Mixture-of-Experts (SG-MoE) to create a scalable library of atomic skills, with a routing encoder for dynamic expert assignment to new skills, facilitating continual learning. Empirical results demonstrate that AtomicVLA outperforms existing VLA models in both simulated (LIBERO, CALVIN) and real-world long-horizon tasks, achieving improvements of up to 21% in real-world continual learning scenarios.

Key Contribution

Forget monolithic action decoders: AtomicVLA's skill-guided mixture-of-experts unlocks significant gains in long-horizon robotic manipulation and continual learning.

Abstract

Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or skills. These challenges present significant barriers for existing VLA models, which use monolithic action decoders trained on aggregated data, resulting in poor scalability. To address these challenges, we propose AtomicVLA, a unified planning-and-execution framework that jointly generates task-level plans, atomic skill abstractions, and fine-grained actions. AtomicVLA constructs a scalable atomic skill library through a Skill-Guided Mixture-of-Experts (SG-MoE), where each expert specializes in mastering generic yet precise atomic skills. Furthermore, we introduce a flexible routing encoder that automatically assigns dedicated atomic experts to new skills, enabling continual learning. We validate our approach through extensive experiments. In simulation, AtomicVLA outperforms $π_{0}$ by 2.4\% on LIBERO, 10\% on LIBERO-LONG, and outperforms $π_{0}$ and $π_{0.5}$ by 0.22 and 0.25 in average task length on CALVIN. Additionally, our AtomicVLA consistently surpasses baselines by 18.3\% and 21\% in real-world long-horizon tasks and continual learning. These results highlight the effectiveness of atomic skill abstraction and dynamic expert composition for long-horizon and lifelong robotic tasks. The project page is \href{https://zhanglk9.github.io/atomicvla-web/}{here}.

Multimodal Models Robotics & Embodied AI Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

Related Papers