Search papers, labs, and topics across Lattice.
OC-VLA++ enhances the OC-VLA framework by integrating geometry-guided paired-view supervision and a cross-view action-equivariance objective to improve viewpoint generalization in robotic manipulation tasks under limited camera coverage. This approach mitigates overfitting to specific viewpoints by ensuring that action predictions remain consistent across different camera angles, rather than relying solely on image-level augmentations. Experimental results reveal significant gains in generalization performance for unseen viewpoints, demonstrating that cross-view action equivariance is a critical factor for robust real-world robotic applications.
Geometry-guided supervision enables robots to generalize actions across viewpoints, achieving robust manipulation even with limited camera coverage.
We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align action supervision with visual observations, camera-space grounding alone can still overfit to the few viewpoints observed during training. OC-VLA++ addresses this limitation by introducing geometry-guided paired-view supervision and an explicit cross-view action-equivariance objective. Given paired observations of the same manipulation scene from geometrically related viewpoints, the model is trained such that their camera-space predictions correspond to the same robot-frame action. This objective explicitly supervises how action predictions should transform across viewpoints, rather than relying solely on image-level augmentation. Experiments demonstrate substantial improvements in unseen-view generalization under limited camera coverage, with performance degrading more gracefully under increasing camera displacement. These results establish cross-view action equivariance as an effective complement to observation-centric action grounding for robust real-world deployment.