Search papers, labs, and topics across Lattice.
This study introduces LongEarth-R1, a vision-language model specifically designed for long-horizon Earth observation reasoning, addressing the limitations of existing models that primarily focus on short sequences. By leveraging a new benchmark, LongEarth-Bench, which includes approximately 120k question-answering samples across diverse tasks, the authors implement a novel training approach that incorporates structured reasoning and group relative policy optimization. The results demonstrate that LongEarth-R1 outperforms existing models on all long-sequence tasks while maintaining competitiveness on standard benchmarks, highlighting its effectiveness in complex spatial and temporal reasoning scenarios.
LongEarth-R1 outperforms all existing models on long-sequence Earth observation tasks, revealing the critical importance of structured reasoning in complex spatial analyses.
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.