Anhui UniversityiFlytekUniversity of Science and TechnologyJun 7, 2026arXiv:2606.08530

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan

AI Summary

The paper introduces GEAR-VLA, a Vision-Language-Action framework that learns geometry-aware action representations to enhance generalization in robotic manipulation across diverse environments and robot embodiments. By employing a coarse-to-fine action learning strategy and integrating a trainable 3D spatial backbone, GEAR-VLA effectively aligns action semantics with continuous action expertise while maintaining embodiment invariance. Experimental results demonstrate significant improvements in performance, achieving state-of-the-art results on multiple benchmarks, including a 90.1% success rate on a universal grasping task with 212 unseen objects.

Key Contribution

GEAR-VLA achieves a remarkable 90.1% success rate in universal grasping tasks, showcasing its ability to generalize across unseen objects and diverse robot embodiments.

Abstract

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from the lack of a unified geometry-aware manipulation representation, leaving existing VLAs vulnerable to low-level trajectory supervision, misaligned 3D features, and embodiment differences. To address this, we propose GEAR-VLA, a VLA framework for learning unified geometry-aware action representations for generalizable robotic manipulation. GEAR-VLA adopts coarse-to-fine action learning, where multi-source embodied pretraining equips the VLM with embodied reasoning and discrete action understanding before latent action tokens connect action semantics to a gradient-decoupled DiT continuous action expert. It further performs semantic-aligned 3D integration by aligning a trainable 3D spatial backbone with the VLA representation while freezing the original VLM-aligned visual pathway. To share this representation across robots, GEAR-VLA uses embodiment canonicalization, where embodiment-aware states and embodiment-invariant actions confine robot differences to the low-level interface. Extensive simulation and real-world experiments demonstrate strong generalization: GEAR-VLA achieves state-of-the-art performance on LIBERO, zero-shot LIBERO-Plus, and RoboTwin 2.0, reaches 85.9% success on AgileX and 81.0% on the pretraining-unseen LDT-01 embodiment, and obtains 90.1% success on a 6,360-trial universal grasping benchmark with 212 unseen objects. Code and models will be released at https://github.com/babynabeauty/GEAR-VLA.

Multimodal Models Robotics & Embodied AI

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Related Papers