PlanBench-V
PlanBench-V evaluates vision-language models on spatial planning map interpretation through an expert-annotated dataset of 223 maps and 1629 question-answer pairs, assessing perception, reasoning, association, and implementation capabilities.
- Released
- 2026-06-04
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing multimodal benchmarks overlook domain-specific spatial planning tasks. PlanBench-V provides a theory-informed framework for evaluating VLM progress in professional planning contexts, identifying persistent limitations in implementation-oriented tasks.
Motivation
Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for decision-making, public communication, and institutional coordination.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.