Search papers, labs, and topics across Lattice.
This paper introduces GeoMEB, a comprehensive multimodal embedding benchmark tailored for urban applications, encompassing 45 diverse tasks such as retrieval and visual question answering, supported by a dataset of 1.32 million examples. It also presents Geo-Embed, a unified embedding model that leverages a shared vision-language backbone for instruction-conditioned query-target matching across various geospatial inputs. The model outperforms existing multimodal embedders by 15.3% on the GeoMEB benchmark, highlighting its effectiveness in handling complex geospatial relationships and temporal changes.
Geo-Embed achieves a 15.3% performance boost over existing models, redefining how we approach multimodal urban understanding.
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. First, we introduce GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, together with training collections comprising 1.32M examples and 286K evaluation queries. Second, we present Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to instruction-conditioned query-target matching over heterogeneous geospatial inputs, including single images, multiple images, text, regions, and masks. On GeoMEB, Geo-Embed achieves the strongest overall performance among representative multimodal embedders, with a 15.3% relative improvement over the strongest baseline. These results motivate future geospatial embedders that organize training and evaluation around explicit query-target relations, including semantic, cross-view, region-level, and temporal correspondence.