Search papers, labs, and topics across Lattice.
This paper introduces ARGUS, a novel observation pre-processing pipeline that utilizes large-scale 3D vision models to align image observations from various camera viewpoints into a canonical viewpoint, enhancing the performance of robot manipulation policies. By addressing the entanglement of scene geometry with viewpoint, ARGUS allows policies to effectively learn from viewpoint-diverse datasets, significantly improving generalization capabilities. Experimental results demonstrate that ARGUS not only outperforms existing methods but also accelerates learning, achieving high success rates 4-6 times faster than previous approaches in both limited-view and view-diverse training scenarios.
Aligning robot scene geometry with ARGUS enables manipulation policies to learn 4-6 times faster from diverse viewpoints, transforming their generalization capabilities.
Large-scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space. This entanglement inherently limits the corresponding policy's ability to learn from viewpoint-diverse datasets (ex. DROID, BridgeV2) and generalize beyond the viewpoints captured in their training data. In this work, we present ARGUS, an observation pre-processing pipeline that uses large-scale 3D vision models to align image observations from arbitrary camera viewpoints into a canonical viewpoint before passing it to downstream visuomotor policies. Experiments across training datasets with varying levels of viewpoint diversity, from fixed multi-view camera configurations to highly varied camera placements, show that our method consistently outperforms prior approaches across both limited-view and view-diverse training regimes. In efficiency comparisons, ARGUS demonstrates an ability to learn from view-diverse data, converging to high success rates 4-6x faster than previous methods by leveraging a simplified observation space. Overall, our findings show that leveraging large-scale 3D vision models reduces the learning burden on visuomotor policies, enabling more efficient learning from large-scale, viewpoint-diverse robot datasets.