MBZUAIWeizmannZayed University of ArtificialJun 11, 2026arXiv:2606.13148

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

Dat Nguyen, Dat Tien Nguyen, Thao Nguyen, Fadillah Adamsyah Maani, F. Maani, Huy M. Le, Muhammad Umer Sheikh, Numan Saeed, Muhammad Haris Khan, Salman Khan

AI Summary

This paper introduces TerraBench, a comprehensive benchmark designed to enhance reasoning over diverse Earth-system data by integrating language models with scientific tools. The benchmark addresses the limitations of existing models that either excel in forecasting or language reasoning but fail to interact effectively with high-dimensional environmental data. Key findings reveal that successful Earth-science agents must not only access tools but also adeptly manage heterogeneous workflows and maintain artifact provenance, as evidenced by the 403 tasks and 24,500 execution steps included in the benchmark.

Key Contribution

Reliable Earth-science agents need to master the orchestration of diverse data sources and workflows, not just tool access.

Abstract

Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real scientific workflows in Earth-science remain underserved. We introduce TerraBench, a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style executable framework that interleaves reasoning, tool calls, and observations to couple LLM planning with scientific tools for environmental retrieval, geospatial processing, simulation, and artifact-backed computation. TerraBench unifies analysis of Earth observation imagery, gridded data, GIS reasoning and simulation in a single executable interface, whereas prior benchmarks isolate these capabilities into narrow individual tasks. It is also the first in this space to pair process-level tool-use metrics with tolerance-aware numeric scoring. The benchmark comprises 403 extensive agentic tasks across three tracks (Fundamentals, Simulator-Grounded, and Document-Grounded Verification) and eight application domains with 24,500 verified execution steps. These results indicate that reliable Earth-science agents must go beyond tool access to coordinate heterogeneous workflows, parameterize tools precisely, and preserve artifact provenance.

Multimodal Models Reasoning & Chain-of-Thought Scientific Discovery & Drug Design

Citation Metrics

Citations0

Influential citations0

References41

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

Related Papers