Benchmark Radar
AI BENCHMARK PROFILE

TerraBench

General AIKnowledge & ReasoningSearch & Retrieval

TerraBench is a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style framework that couples LLM planning with scientific tools. It includes 403 tasks across three tracks and eight domains with 24,500 verified steps.

Released
2026-06-11
Readiness
Paper only
Primary field
General AI

Why it matters

Earth-science workflows require reasoning over heterogeneous data types. TerraBench unifies these capabilities in a single interface, with process-level metrics and tolerance-aware scoring.

Motivation

Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.