TerraLogic
TerraLogic evaluates hierarchical geospatial reasoning in Earth observation through 545 scenario-driven tasks spanning optical, SAR, and infrared imagery. Tasks include hazard vulnerability assessment, urban heat island analysis, and forest fragmentation dynamics. Evaluation uses tool-augmented agents with verifiable multi-step workflows and scored by step-wise and end-to-end metrics.
- Released
- 2026-07-14
- Readiness
- Runnable
- Primary field
- Cybersecurity
Why it matters
Existing remote sensing benchmarks primarily target perception tasks, leaving a gap in assessing cognitive-level geospatial reasoning. TerraLogic provides a fixed dataset and protocol for comparing agent performance on compositional, long-horizon analysis, enabling systematic evaluation of tool-augmented reasoning across modalities.
Motivation
Beyond perception, reasoning is essential in remote sensing for advanced interpretation, inference, and decision-making.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.