Search papers, labs, and topics across Lattice.
This study investigates the effectiveness of behavior-aligned representations, such as object bounding boxes and end-effector traces, in facilitating cross-embodiment transfer within vision-language-action models for robot manipulation. By developing a simulation-based benchmark, the authors demonstrate that these representations can significantly enhance the transfer of learned behaviors across different robot embodiments, particularly when larger datasets are utilized. Notably, the introduction of end-effector traces leads to a 28% improvement in task completion for real robot policies trained on simulation data, underscoring the potential for improved practical applications in robotics.
End-effector traces boost cross-embodiment transfer in robot manipulation, enhancing real-world task performance by 28% when leveraging simulation data.
Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned representations (e.g., object bounding boxes, language motions, end-effector traces of robot motion) in vision-language-action (VLA) models to promote cross-embodiment transfer. We hypothesize that by possessing invariances across embodiments while being predictive of robot actions, these representations can help unify large-scale cross-embodiment data to enhance transfer. To assess our hypothesis, we develop a simulation-based benchmark designed to assess transfer with diverse cross-embodiment data to new embodiments. Using this benchmark, we compare different representations and ways of incorporating them. We identify that end-effector traces can be particularly beneficial for transfer, representations are generally more useful with larger prior datasets, and can be used to benefit from action-free data. We also demonstrate that they can enhance sim-to-real cross-embodiment transfer, improving task completion progress of real robot policies pre-trained on simulation data by 28%. We provide videos of our evaluations at our website: https://ajaysridhar.com/barx/.