Search papers, labs, and topics across Lattice.
RoadVGGT introduces a novel feed-forward framework for road surface reconstruction that leverages multi-view images and geometric foundation models to predict compact Gaussian representations without requiring per-scene optimization. This approach addresses the limitations of existing methods, which are constrained by scene-dependent training and coverage design, thereby enabling scalable reconstruction of newly collected roads. The framework not only enhances the quality of RGB and semantic maps but also improves elevation estimation and novel view synthesis, showcasing significant advancements in road surface representation.
Eliminating per-scene optimization, RoadVGGT achieves scalable and high-quality road surface reconstruction using a compact Gaussian representation.
Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road--sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird's-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.