Search papers, labs, and topics across Lattice.
The paper introduces NebulaVLA, an asynchronous dual-frequency Vision-Language-Action model that separates high-level semantic reasoning from low-level action control, addressing efficiency-performance trade-offs in robotic manipulation. By employing a unified language-grounded action representation called GESTURE-7 and a Guide Action algorithm that ensures kinematic continuity, the model enhances cross-embodiment generalization and execution smoothness. Evaluations reveal that NebulaVLA achieves an 85.5% success rate on the LIBERO-Plus benchmark while accelerating action generation by approximately 2.7 times compared to synchronous models.
Achieving an 85.5% success rate in robotic manipulation while accelerating action generation by 2.7 times could redefine efficiency in VLA systems.
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.