Search papers, labs, and topics across Lattice.
This paper introduces AdaRoboVLG, a novel Vision-Language-Grasp (VLG) framework that decouples generalizable grasp synthesis from task-specific understanding, allowing for efficient learning across different robotic hands. By employing a base policy that generates and evaluates grasp candidates through kinematic mapping and stability estimation, the framework integrates composable priors from specialized foundation models to enhance adaptability without retraining. Experimental results reveal that AdaRoboVLG achieves strong cross-hand generalization and maintains performance in complex environments, marking a significant advancement in scalable robotic grasping methodologies.
Efficiently decoupling grasp synthesis from task understanding allows AdaRoboVLG to generalize across robotic hands while maintaining high performance in dynamic environments.
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/