Search papers, labs, and topics across Lattice.
This paper introduces GuidedAttention, a visuomotor imitation learning framework that enhances robot manipulation by incorporating interpretable and correctable visual attention as an intermediate representation. By predicting task-relevant attention keypoints from camera images and allowing users to correct these keypoints at the start of the rollout, the framework ensures that the corrected attention is maintained throughout the execution via a tracking module. Experimental results show that GuidedAttention significantly improves performance in both simulation and real-world scenarios, particularly when faced with out-of-distribution positional and appearance challenges.
Correctable visual attention can drastically enhance robot manipulation performance, especially in unpredictable environments.
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions.