Benchmark Radar
AI BENCHMARK PROFILE

GISAgentBench

General AIKnowledge & Reasoning

GISAgentBench evaluates LLM agents on multi-step GIS tasks from practitioner sources, with 349 tasks and executable reference trajectories.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Provides deterministic evaluation with ground truth outputs, addressing limitations of surrogate signals.

Motivation

Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.