Search papers, labs, and topics across Lattice.
This paper enhances 3D Gaussian splatting (3DGS) for novel view synthesis by integrating multi-view geometric priors, specifically predicted normal and depth maps, to address the limitations of traditional structure-from-motion initialization and photometric optimization. The authors demonstrate that utilizing multi-view predictions from the visual geometry grounded transformer (VGGT) leads to superior reconstruction quality compared to single-view methods, particularly in challenging scenarios involving high-specularity objects. Their extensive experiments reveal that incorporating a confidence map derived from multi-view models significantly boosts the effectiveness of these priors, resulting in consistent improvements across standard benchmarks.
Multi-view geometric priors can drastically enhance 3D reconstruction quality, especially for complex scenes with specular surfaces.
3D Gaussian splatting (3DGS) has emerged as a widely-used tool for novel view synthesis, offering real-time rendering in a sparse representation. However, the method's reliance on structure-from-motion initialization and photometric optimization can lead to suboptimal geometric reconstruction, particularly for objects with high specularity. In this work, we investigate the integration of geometric priors, in the form of predicted normal and depth maps, into the 3DGS framework to improve the reconstruction quality. We analyze the effect of incorporating these priors into GS-based methods and our evaluation reveals that multi-view predictions, as they are done by the recent visual geometry grounded transformer (VGGT), outperform single-view alternatives. A major factor is the existence of a confidence map for the estimations, which comes as a by-product of multi-view models and which can significantly improve the effectiveness of priors by weighting each prediction appropriately. Extensive experiments on standard benchmarks show consistent improvement in reconstruction quality and significant gains in complex scenes including specular objects.