MultiGlobeQA
MultiGlobeQA evaluates geospatial reasoning in large language models across 46,060 question-answer pairs in 17 languages, covering 14 spatial-function families and 15 answer formats, with ground truth from three knowledge graphs.
- Released
- 2026-08-04
- Readiness
- Paper only
- Primary field
- Transport & Logistics
Why it matters
Current geospatial benchmarks are limited in language coverage and geographic control. This benchmark aims to provide a broader, multilingual evaluation to identify specific failures in computational reasoning over geographic knowledge.
Motivation
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.