Search papers, labs, and topics across Lattice.
This paper introduces a unified framework for enhancing geometry-consistent and domain-robust image representations in monocular endoscopy, addressing challenges such as limited depth cues and substantial appearance variation. By integrating a synthetic data pipeline with Hierarchy-Aware Geometry-Semantic Adaptation, the method improves feature correspondence and semantic consistency, which are critical for reliable navigation tasks. Experimental results demonstrate significant advancements in pose estimation and monocular depth estimation, with effective transfer from synthetic to real clinical scenarios.
Geometry-guided adaptation can dramatically enhance the reliability of endoscopic navigation, overcoming traditional challenges in depth perception and feature alignment.
Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment. Although recent vision foundation models have shown promise, their learned representations often remain insufficiently geometry-consistent, hindering stable feature correspondence and limiting their reliability for downstream navigation tasks. We propose a unified framework for learning geometry-consistent and domain-robust image representations for monocular endoscopy. The framework combines a synthetic data pipeline that provides accurate geometric supervision with Hierarchy-Aware Geometry-Semantic Adaptation, a structured alternative to standard LoRA that inserts low-rank adapters selectively across the transformer hierarchy and couples them with layer-wise training objectives to encourage geometric correspondence in intermediate features and semantic consistency in deeper features. Experiments on public and proprietary datasets show improved geometric and semantic representation quality, leading to better performance on downstream navigation tasks including pose estimation and monocular depth estimation. The learned representations show favorable synthetic-to-real transfer on clinical bronchoscopy and provide a useful initialization for adaptation to sinus endoscopy and colonoscopy under limited supervision. The framework also shows favorable scaling with model size and training data. These results support hierarchy-aware, geometry-guided adaptation as a practical approach for endoscopic representation learning.