Search papers, labs, and topics across Lattice.
This paper introduces a joint-view consensus-guided learning framework for cross-view geo-localization that eliminates the need for geometric warping, which often leads to visual distortions and fragile correspondences. By dynamically mining and strengthening semantic consensus directly in the feature space, the method allows for effective cross-view interaction and robust retrieval. Extensive experiments show that this approach achieves state-of-the-art performance on four standard benchmarks, highlighting the significance of semantic consensus in reliable geo-localization tasks.
Bypassing geometric warping, this framework reveals that mining semantic consensus directly in feature space can dramatically enhance cross-view geo-localization accuracy.
Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.