Search papers, labs, and topics across Lattice.
This paper introduces the Geometry-grounded Unified 3D Perception (GeoUP) framework, which integrates metric 3D structure into camera-based autonomous driving perception by adapting the latent space of VGGT for synchronized multi-camera streams. By employing a factorization approach that incorporates self, temporal, and view attention, GeoUP captures distinct correspondences while utilizing calibration-aware raymap encodings to ensure accurate metric scale and camera geometry. Extensive evaluations across multiple datasets, including nuScenes and Waymo, reveal that GeoUP achieves state-of-the-art performance in depth estimation, 3D object detection, and semantic occupancy prediction, underscoring the importance of geometry-grounded representations in autonomous driving applications.
Achieving state-of-the-art performance in 3D perception tasks, GeoUP reveals that integrating geometry-grounded representations can significantly enhance autonomous driving systems.
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific modules. As a result, their shared representations may fail to preserve explicit metric geometry and consistent 3D scene structure. In this paper, we present a Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes. GeoUP factorizes cross-image interaction into self, temporal, and view attention to capture structurally distinct temporal and cross-view correspondences. It further injects calibration-aware raymap encodings to provide metric scale and camera geometry. The resulting geometry-grounded latent is decoded for metric depth estimation, 3D object detection, and semantic occupancy prediction, corresponding to surface-, instance-, and volume-level readouts of the same 3D scene. Through joint multi-task and multi-dataset training, GeoUP effectively leverages heterogeneous annotations and generalizes across diverse sensor configurations and perception ranges. Extensive experiments on nuScenes, Argoverse 2, Waymo, KITTI, and DDAD demonstrate that GeoUP achieves SOTA performance across detection, occupancy, and depth estimation. These results validate the effectiveness of geometry-grounded representations for unified 3D driving perception.