Search papers, labs, and topics across Lattice.
This paper analyzes the uncertainty quality of the Visual Geometry Grounded Transformer (VGGT) on the DTU Benchmark Dataset, focusing on its ability to provide reliable uncertainty estimates alongside high reconstruction accuracy. The study identifies an effective confidence threshold that can filter VGGT's raw outputs, thereby enhancing the accuracy of its 3D reconstructions. These findings underscore the importance of uncertainty quantification in photogrammetry, paving the way for more trustworthy and robust 3D reconstruction methods.
Enhancing uncertainty quality in VGGT can significantly boost the accuracy of 3D reconstructions, making real-time photogrammetry more reliable.
Visual Geometry Grounded Transformer (VGGT) has already attracted a great deal of attention in a short period of time, not least due to the Best Paper Award at CVPR-2025. Similar to DUSt3R and MASt3R, VGGT aims to bring about a paradigm shift by replacing established methods like bundle adjustment and feature matching with a simple, unified, feed-forward neural network that predicts camera poses, depth maps, and dense 3D structure directly from multiple images of a scene in a few seconds. A key aspect is its ability to process an arbitrary number of views consistently in a single forward pass without any post-processing or iterative optimization. For photogrammetry, this opens new possibilities for real-time, scalable, and accessible 3D reconstruction. In this context, not only high reconstruction accuracy but also high-quality uncertainty estimates are crucial, as they foster trust and enable robust quality assurance. This paper therefore investigates the quality of VGGT's uncertainty predictions. The analysis identifies an effective confidence threshold for filtering VGGT's raw output and demonstrates that enhancing uncertainty quality holds strong potential for improving the accuracy of its 3D reconstructions.