Search papers, labs, and topics across Lattice.
This paper introduces RA-VLA, a retrieval-augmented Vision-Language-Action framework designed to enhance test-time adaptation in robotic manipulation tasks. By integrating behavior-aligned context retrieval with a grounded execution pipeline, RA-VLA overcomes the adaptation bottleneck seen in existing In-Context Imitation Learning methods, which struggle with translating expert context into executable actions due to superficial retrieval and behavioral inertia. Empirical evaluations reveal that RA-VLA significantly improves success rates and computational efficiency on the LIBERO benchmark and in real-world scenarios, marking a substantial advancement in training-free robotic adaptation.
RA-VLA transforms robotic adaptation by seamlessly integrating context retrieval with execution, achieving superior performance without the need for training.
Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions. This failure originates from superficial retrieval mechanisms and an inherent behavioral inertia that anchors the policy to its pre-trained priors. To address these limitations, we present RA-VLA, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline. By enforcing faithful adherence to functional cues within a scalable architecture, RA-VLA facilitates seamless task adaptation while preserving inference efficiency. Our empirical evaluations across the LIBERO benchmark and a real-world UR5e environment demonstrate that RA-VLA achieves superior success rates and computational efficiency, establishing a robust framework for training-free robotic adaptation.