Search papers, labs, and topics across Lattice.
This paper introduces EDAR, an Environment-Dependent Action Representation that integrates executable control structures with expected visual outcomes to enhance robotic manipulation. By addressing the inherent variability of action semantics based on environmental context, EDAR enables robots to learn more effective action representations that capture the nuances of interaction semantics. Experimental results show significant improvements in policy learning for both simulated and real-robot manipulation tasks, particularly in long-horizon scenarios.
Grounding action representations in environmental context can drastically improve robotic manipulation performance, especially for complex tasks.
Learning effective action representations is critical for robotic manipulation, where raw control trajectories are often noisy, redundant, and difficult to model directly. Existing methods mainly encode the structure of the action stream itself, treating the role of actions in the environment as implicit. Yet manipulation is about changing the world: the same action segment can induce different outcomes under different scene contexts, making action semantics inherently environment-dependent. We propose EDAR, an Environment-Dependent Action Representation that grounds action tokens in both executable control structure and expected visual consequences. By coupling motor commands with their environment-conditioned effects, EDAR encourages the learned action space to capture interaction semantics rather than merely command-level patterns. Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation. These results highlight the importance of grounding action representations in executable control structure and environment-conditioned visual change.