Microsoft ResearchUMDMar 2, 2026arXiv:2603.02080

From Pixels to Patches: Pooling Strategies for Earth Embeddings

Isaac Corley, Caleb Robinson, Inbal Becker-Reshef, Juan M. Lavista Ferres

AI Summary

This paper investigates pooling strategies for aggregating pixel-level embeddings into patch representations for geospatial foundation models, addressing the limitations of mean pooling in preserving class-discriminative information and handling spatial shifts. The authors introduce EuroSAT-Embed, a dataset of 81,000 embedding GeoTIFFs derived from AlphaEarth, OlmoEarth, and Tessera, to benchmark 11 training-free and 2 parametric pooling methods. Results demonstrate that richer pooling schemes like GeM and Stats pooling significantly improve geographic generalization and accuracy compared to mean pooling, with GeM recommended as a simple drop-in replacement.

Key Contribution

Ditch mean pooling in your geospatial foundation models: richer pooling methods like GeM can boost accuracy by up to 5% and slash the geographic generalization gap by 40%.

Abstract

As geospatial foundation models shift from patch-level to pixel-level embeddings, practitioners must aggregate thousands of pixel vectors into patch representations that preserve class-discriminative signal while matching downstream label resolution. The default choice, mean pooling, discards within-patch variability and can drop accuracy by more than 10% under spatial shift. To evaluate this effect, we introduce EuroSAT-Embed: 81,000 embedding GeoTIFFs derived from three foundation models: AlphaEarth, OlmoEarth, and Tessera. We benchmark 11 training-free and 2 parametric pooling methods under both random and geographically disjoint test splits. Our results show that richer pooling schemes reduce the geographic generalization gap by up to 40% relative to mean pooling and increases accuracy by up to 5% on spatial splits. We recommend Generalized Mean Pooling (GeM) as a drop-in replacement for mean pooling: it improves accuracy without increasing embedding dimensionality. For maximum accuracy, Stats pooling (concatenation of min/max/mean/std pooling) performs best at 4x the embedding size. We further find that pooling effectiveness varies across embedding sources and that higher-dimensional embeddings benefit most from distributional statistics.

Computer Vision Data Curation & Synthetic Data Multimodal Models

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

From Pixels to Patches: Pooling Strategies for Earth Embeddings

Related Papers