Search papers, labs, and topics across Lattice.
This paper introduces a Concept Expert module designed to enhance Vision-Language-Action (VLA) models by integrating 3D structural information and commonsense knowledge, which are often overlooked in traditional 2D inputs. By employing a two-phase approach鈥攊nitially estimating kinematic parameters from Vision Foundation Models and subsequently tracking dynamic concept parameters during manipulation鈥攖he method enables VLA models to execute complex, high-precision tasks more effectively. Experimental results reveal significant improvements in both success rates and learning efficiency, highlighting the advantages of structured, analytic guidance in VLA post-training.
VLA models can achieve higher precision and adaptability by leveraging 3D structural insights, leading to improved manipulation success rates.
Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we construct a Concept Expert module for VLA to build executable Analytic Concepts that represent objects as explicit, programmatic blueprints. Our mechanism operates in two synergistic phases: First, prior to VLA inference, the Concept Expert leverages 3D information from Vision Foundation Models (VFMs) to estimate the initial kinematic and structural parameters. Second, throughout the manipulation process, the VLA model utilizes its inherent capability to dynamically track the dynamic concept parameters, continuously aligning them with observational changes to ensure persistent accuracy. Once established, the Analytic Concepts provide explicit, high-quality guidance for VLA fine-tuning through (1) dense, programmatic manipulation rewards and (2) precise spatial guidance. This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility. Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.