Search papers, labs, and topics across Lattice.
This paper introduces TeaMatch, a framework for enhancing cross-modal representation learning in 2D-3D matching by leveraging teachability as a key criterion. The method addresses the challenges of maintaining reliable correspondences in the presence of noisy inputs and ambiguous structures by training weak learners to recover representations under degraded conditions. Experimental results show that TeaMatch significantly improves robustness and achieves state-of-the-art performance on various benchmarks, demonstrating its effectiveness in real-world applications.
Teachability in representation learning boosts 2D-3D matching robustness, achieving state-of-the-art results even in challenging conditions.
Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods, existing approaches primarily optimize matching within a single model and often struggle to maintain reliable correspondences under challenging conditions such as noisy inputs, low overlap, and ambiguous structures. In this work, we propose TeaMatch, a novel framework that introduces teachability as a criterion for cross-modal representation learning. We define teachability as the ability of a representation to be effectively recovered by weak learners under degraded inputs, reflecting its structural consistency and robustness. To this end, we construct a set of task-specific weak students that simulate common failure modes and train them to imitate the teacher on a training split while evaluating their recoverability on a disjoint meta split. The teacher is then optimized to improve the students' ability to recover reliable correspondences, guided by correspondence-level and geometry-aware constraints. Our framework can be seamlessly integrated into existing coarse-to-fine matching pipelines without additional inference cost. Extensive experiments demonstrate that TeaMatch improves matching robustness and achieves state-of-the-art performance on challenging 2D-3D matching benchmarks.