Search papers, labs, and topics across Lattice.
This survey presents a unified framework for robot learning by integrating representation learning, vision-language-action (VLA) models, and world models, addressing the fragmentation that hampers generalization and long-horizon reasoning in robotic systems. By categorizing existing methods and analyzing their interactions, the authors identify key challenges such as uncertainty quantification and cross-embodiment transfer that hinder effective deployment in complex environments. The paper outlines future directions for creating robust, integrated robotic systems that can maintain consistent internal representations and support decision-making over extended interactions.
Fragmentation in robot learning systems limits their effectiveness, but a unified framework could enable robots to reason and act more reliably in complex environments.
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enabling robots to work in increasingly complex environments. However, these paradigms are typically developed in isolation, resulting in fragmented systems that struggle with generalization, long-horizon temporal reasoning and planning, and deployment in unstructured environments. In this survey, we present a unified perspective on robot learning by organizing the existing methods along three complementary axes: understanding through representation learning, acting through VLA models, and reasoning through world models. We introduce a structured taxonomy that captures key design choices in environment representation, policy learning, and predictive modeling, and summarize the recent progress in these domains. Beyond classifying the existing works, we analyze how these components interact, discuss common limitations, and highlight emerging trends towards more integrated systems. Through this lens, we identify the challenges in the domain of robot learning, including uncertainty quantification, out-of-distribution generalization, cross-embodiment transfer, long-context understanding, and long-horizon planning. We argue that these challenges arise not only from limitations within individual components but also from the lack of integration across perception, action, and reasoning. Building on this analysis, we outline future directions towards unified, physically grounded, and probabilistic robot learning to develop robust robotic systems that maintain consistent internal representations and support decision making over extended interactions in real-world environments.