Search papers, labs, and topics across Lattice.
This study systematically compares two Geospatial Foundation Models (GFMs), THOR and TerraMind, to elucidate the factors contributing to their performance differences, focusing on architectural design choices, decoder complexity, and use-case-specific characteristics. By analyzing ten diverse use cases, the authors reveal that architectural elements, particularly patch size and decoder type, account for more variance in performance than the models themselves. The findings suggest a nuanced understanding of model performance, advocating for a diagnostic ablation methodology that can inform future GFM development beyond the specific models studied.
Architectural choices, not model identity, drive performance differences in Geospatial Foundation Models, challenging conventional ranking methods.
Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact? This study addresses that gap through a controlled comparison of two GFMs developed under European Space Agency's $\Phi$-lab with contrasting design philosophies: THOR, which introduces a compute-adaptive architecture supporting variable patch sizes and unifies Sentinel-1, -2, and -3 data at their native resolutions; and TerraMind, a multimodal generative GFM pretrained with a dual-scale token/pixel objective that enables any-to-any cross-modal generation (Thinking-in-Modalities) to infer missing sensors at inference time. Rather than reporting a single leaderboard, we investigate the axes along which the two architectures actually differ - patch size, decoder complexity, finetuning regime, input modality, and model scale - across ten use cases spanning segmentation and regression in diverse domains, including climate disaster response, methane leak detection, snow monitoring, or sea ice mapping. We find that architectural design choices - patch size and decoder type in particular - explain more performance variance than model identity itself, that the two models embody complementary investment strategies (pretraining-time scale for TerraMind versus inference-time tokenisation for THOR), and that correctly interpreting results requires dataset-level characterisation. The resulting picture is not a single winner but a set of hypotheses and a diagnostic ablation methodology that we expect to generalise to future GFMs beyond THOR and TerraMind.