Search papers, labs, and topics across Lattice.
This survey comprehensively reviews the evolution of monocular depth estimation, detailing the transition from early learning-based methods to the current landscape dominated by foundation models. It categorizes depth estimation approaches into relative and metric frameworks, while also analyzing the impact of large-scale pretraining and synthetic data on model performance. Key findings reveal that foundation models significantly enhance accuracy and robustness, paving the way for advanced applications in robotics and augmented reality.
Foundation models are revolutionizing monocular depth estimation, achieving unprecedented accuracy and robustness in real-world applications.
Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative foundation models. We begin by framing the problem, distinguishing between relative and metric depth estimation, and highlighting the key challenges that have shaped a decade of research. We then present common problem formulations and introduce the most widely used datasets, covering indoor, outdoor, and synthetic data. Following this, we review major advances prior to the foundation model era, distilling core insights from influential methods that contributed to improvements in accuracy, efficiency, and robustness. The survey then turns to the recent surge of foundation-model-based approaches, categorizing them into discriminative and generative paradigms and emphasizing the critical roles of large-scale pretraining (e.g., DINOv3) and synthetic data. We compare representative models using both quantitative benchmarks and qualitative examples, and discuss natural extensions to video-based depth estimation. Further, to illustrate real-world impact, we highlight the integration of depth estimation into applications such as visual SLAM, content generation, and robot perception. Finally, we outline open challenges and promising research directions as the field advances further into the era of foundation models.