AI BENCHMARK PROFILE
GISAgentBench
GISAgentBench evaluates LLM agents on multi-step GIS tasks from practitioner sources, with 349 tasks and executable reference trajectories.
- Released
- 2026-08-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides deterministic evaluation with ground truth outputs, addressing limitations of surrogate signals.
Motivation
Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.