Search papers, labs, and topics across Lattice.
This paper introduces Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration (DPA-I2P), which enhances camera pose estimation by addressing the challenges of cross-modal correspondence between images and sparse LiDAR point clouds. The method employs Ray-Conditioned Metric Depth Encoding (RMDE) and Projection-Consistent Vision Lifting (PVL) to leverage depth and visual cues in a geometry-aware manner, alongside Cross-Modal Query Pruning (CQP) to improve matching stability. Experimental results on the KITTI and nuScenes datasets show significant improvements, with DPA-I2P reducing Relative Translation Error (RTE) and Relative Rotation Error (RRE) by 45.0% and 55.6% respectively compared to the strongest baseline.
DPA-I2P achieves a remarkable 45% reduction in pose estimation errors, setting a new standard for image-to-point cloud registration in autonomous driving.
Image-to-Point Cloud Registration aims to estimate the camera pose of a given image within a 3D scene point cloud, which is a fundamental task in autonomous driving and large-scale outdoor localization. Recent implicit correspondence learning methods have improved registration performance by learning cross-modal alignment in an end-to-end framework, leading to more accurate camera pose estimation. However, due to the inherent modality discrepancy between images and sparse LiDAR point clouds, reliable cross-modal correspondence learning remains challenging. To address this issue, we propose Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration (DPA-I2P). Unlike naive depth or feature concatenation, Ray-Conditioned Metric Depth Encoding (RMDE) and Projection-Consistent Vision Lifting (PVL) exploit depth and visual cues in a structured, geometry-aware manner. In addition, Cross-Modal Query Pruning (CQP) suppresses unreliable queries during early refinement to improve matching stability. Experiments on KITTI and nuScenes demonstrate the effectiveness of the proposed method. On KITTI, DPA-I2P reduces RTE and RRE by 45.0% and 55.6% over the strongest implicit baseline, respectively. On nuScenes, DPA-I2P also improves registration accuracy over the evaluated baselines, suggesting better transferability to different driving scenes.