Search papers, labs, and topics across Lattice.
This paper introduces VIDAR, a visual-inertial dense reconstruction framework that integrates SVO+IMU odometry with the Depth Anything 3 foundation model to achieve stable metric scale in monocular dense geometry. By using the visual-inertial front end as a metric anchor, VIDAR effectively aligns dense predictions over time, resulting in significant improvements in scale accuracy and reconstruction quality. Experimental results demonstrate that pose injection reduces scale error to approximately 1% and achieves a mean F@0.10 score of 0.463, while a decoupled alignment strategy further enhances performance to 0.676 without relying on ground-truth poses.
Achieving a scale error reduction to about 1% in monocular dense reconstruction could redefine the standards for visual-inertial systems.
Monocular foundation models provide dense geometry but usually lack a stable metric scale. This paper presents VIDAR, a visual-inertial dense reconstruction framework that couples SVO+IMU odometry with Depth Anything 3. VIDAR uses the visual-inertial front end as a metric anchor: it provides camera poses, scale, and a consistent world frame for aligning dense foundation-model predictions across time. The foundation model then contributes detailed local geometry that is fused into a global reconstruction. We study both pose-conditioned DA3 and a decoupled alignment strategy. On EuRoC, pose injection reduces scale error to about 1\% and reaches 0.463 mean F@0.10; the decoupled hybrid improves this to 0.676 without ground-truth poses. Results on EuRoC and TUM RGB-D show that VIDAR is a practical route to metric dense monocular reconstruction.