Benchmark Radar
AI BENCHMARK PROFILE

MapReason-OSM

Transport & LogisticsMultimodal PerceptionVi-Sri

MapReason-OSM evaluates vision-language models on graph-verifiable mobility decisions from self-rendered OpenStreetMap panels. It covers 12 tasks in routing, facility location, and visual disambiguation, with structured outputs scored against hidden oracles for validity, legality, optimality, and constraint satisfaction, plus cross-zoom consistency.

Released
2026-06-21
Readiness
Runnable
Primary field
Transport & Logistics

Why it matters

Existing map benchmarks often rely on free-text or multiple-choice answers that cannot be verified against road networks. This benchmark provides a reproducible, exact scoring contract for decision-making tasks, supporting comparison of VLM capabilities in logistics, delivery, and accessible navigation.

Motivation

Vision-language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable decision (a route, a pin, a parking choice) that must respect the road network.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.