Search papers, labs, and topics across Lattice.
This paper addresses a significant limitation in vision-language-action (VLA) models, where the action expert lacks sufficient access to critical 3D geometric and 2D semantic information from visual language models (VLM). The authors introduce V-Link, which enhances feature transfer from VLM to action by incorporating Spatial and Semantic Query representations, leading to improved perceptual grounding and performance in robotic manipulation tasks. Experimental results demonstrate that V-Link achieves notable increases in success rates across multiple benchmarks, including a +31.2% improvement on LIBERO-Plus and +24% on real-world humanoid tasks.
V-Link bridges the gap between visual perception and action control, boosting robotic manipulation success rates by up to 31.2% through enhanced feature integration.
Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.