VertiCue-Bench
VertiCue-Bench evaluates multimodal large language models on geospatial reasoning using canopy height models to resolve 2D ambiguity in remote sensing natural scenes. It includes 1,534 instances across 17 tasks, testing height perception and semantic reasoning.
- Released
- 2026-05-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current remote sensing benchmarks are mostly 2D-centric, failing in environments with spectral confusion. This benchmark addresses the gap of whether models can leverage vertical cues for semantic disambiguation, providing insights into geometry-to-semantics reasoning in MLLMs.
Motivation
Multimodal Large Language Models (MLLMs) have recently shown promising progress in geospatial reasoning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.