Search papers, labs, and topics across Lattice.
This paper investigates the "embodiment gap" in robot foundation models (RFMs), which refers to the disparity between reusable model components and their practical execution on specific robotic platforms. The authors categorize existing methods based on their adaptability and the type of shared structure, providing a comprehensive framework to assess the necessary adaptations for effective deployment. Key findings reveal that while certain aspects of RFMs can be generalized, significant work remains to tailor these models to individual robotic embodiments, impacting their practical usability in real-world applications.
The embodiment gap reveals that even advanced robot foundation models require substantial adaptation to function on specific robotic platforms, challenging the assumption that scaling alone suffices for generalization.
Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should improve generalization. In robotics, however, a model can generalize while work still remains before it can run on a robot with a particular body. The work required differs across methods and target robots, and those differences affect practical deployment. We call the gap between reusable models, representations, or data and their use in execution on the target robot the embodiment gap. This survey examines what can be reused across robot embodiments and what must still be implemented on a new robot. We place existing methods on a two-axis map that shows the type of shared structure and the stage at which adaptation is needed for execution on the target robot. We then examine recent work through three overlapping research directions: sharing semantics and perception, sharing robot data and interfaces, and learning correspondence across embodiments. We also propose a reporting framework for adaptation work that success rate alone does not reveal. The framework identifies the work that should be checked when comparing cross-embodiment learning and highlights work that remains on a new robot and questions for future study.