Search papers, labs, and topics across Lattice.
This paper introduces GIFT (Geometry-Invariant Fine-Tuning), a novel framework designed to enhance monocular depth estimation models' performance on non-Lambertian surfaces without requiring depth labels. By leveraging the geometric invariance of surfaces despite varying appearances, GIFT effectively reduces depth hallucinations in challenging scenarios like mirrors and glass while maintaining overall depth estimation accuracy. Experimental results on a new benchmark and a real-world dataset reveal that GIFT significantly improves depth prediction in complex environments, offering a cost-effective solution for real-world applications.
GIFT reduces depth hallucinations in non-Lambertian surfaces without needing depth labels, enhancing monocular depth models' practical utility.
Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirrors or transmitted content behind glass rather than the physical surface itself. Adapting these models with real-world data is challenging because conventional depth sensors are also unreliable in such regions. We observe that while the appearance of a non-Lambertian surface varies with its reflected or transmitted environment, its underlying geometry remains unchanged. Based on this observation, we propose GIFT (Geometry-Invariant Fine-Tuning), a parameter-efficient post-training framework that requires no measured depth labels. We collect groups of RGB images under controlled appearance changes while keeping the camera and target geometry fixed. GIFT exploits geometric invariance across these observations to suppress non-Lambertian depth hallucinations while retaining general depth estimation capability. We further construct a controlled benchmark that evaluates non-Lambertian depth recovery, robustness to appearance changes, and performance retention in other regions. Experiments on our benchmark and an independent real-world dataset demonstrate that GIFT improves depth prediction for mirrors and transparent objects while largely preserving the base model's performance, providing a practical and low-cost approach for adapting monocular depth foundation models to non-Lambertian scenes.