DRIVESPATIAL
DriveSpatial is a benchmark of 15.6K human-verified QA pairs across 20 tasks from five AD datasets, evaluating VLMs on cognitive scene construction, multi-view relational understanding, temporal reasoning, and generalization.
- Released
- 2026-05-22
- Readiness
- Paper only
- Primary field
- Transport & Logistics
Why it matters
Existing AD vision-language benchmarks focus on static, single-view QA, leaving unclear whether VLMs can reason over dynamic driving scenes. DriveSpatial provides a multi-sourced, human-verified protocol to measure spatiotemporal reasoning and reveals a substantial human-model gap.
Motivation
Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations, interactions, and future dynamics.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.