Search papers, labs, and topics across Lattice.
×10−42\times 10^{-4} decaying according to a cosine annealing schedule to a terminal learning rate of 10−610^{-6}. We sample 512 keypoints per image with an NMS radius of 3px following the training sampling of [20]. For the detector we set dmax=d_{\text{max}}{=}1.2px, ρpos=\rho_{\text{pos}}{=}1, and ρneg=−min{10−2,t⋅10−6}\rho_{\text{neg}}{=}-\min\{10^{-2},t\cdot 10^{-6}\} where tt is the number of optimizer steps taken through training. At inference time we use subpixel sampling based on the soft-argmax over the patch around the selected keypoint [20, 72]. For the training of the other two modules, the inference setting of the detector is used to capture the distribution of keypoints at inference. The covariance estimator head is trained for 20k steps and the ranker module is separately trained for 1 epoch with the identical setup as the detector training. For both training runs we freeze all model layers except those of the pertinent module/head. More details are provided in the supplementary. 4.2 Two-View Keypoint Matching Setup: We present a comprehensive evaluation of our keypoint detector in the two-view setting. We use 4 different evaluation datasets and use the ground truth transformations to project keypoints across views. Keypoint matching is subsequently performed by identifying mutual nearest neighbors within a specified reprojection radius in both views. We force all detectors to detect the same number of keypoints. HPatches [4] comprises over 500 real-world image pairs under a homography transformation. The views are subject to either illumination or viewpoint changes. DNIM [73] consists of 1722 images grouped into 17 sequences per webcam. This data set exhibits strong illumination changes, and we sample 428 random image pairs augmented with random homographies to also evaluate the robustness to perspective changes. MegaDepth [36] is a large-scale dataset of photo-tourism internet images. We use the subset MegaDepth1800 [37] of 4 scenes from the dataset’s test set. We introduce the ETH
3
0
6
12
VidMap achieves unprecedented robustness and accuracy in metric reconstruction from uncalibrated videos, outperforming both SLAM and SfM methods under extreme conditions.
Unlock zero-shot geospatial reasoning by jointly embedding satellite imagery, street view, elevation, text, and lat/lon coordinates into a single space.
Forget complex architectures: RaCo achieves SOTA keypoint matching and repeatability by cleverly combining ranking and covariance estimation in a lightweight network, trained without covisible image pairs.