AirGroundBench
AirGroundBench evaluates multi-view spatial intelligence in multimodal large language models through UAV-UGV collaborative tasks, including 62,000 dual-view multiple-choice questions and 115 navigation episodes across 11 simulated environments.
- Released
- 2026-06-26
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
This benchmark addresses the gap in assessing geometric consistency across heterogeneous views, providing a structured evaluation for capabilities like cross-view alignment and spatial reasoning that are critical for embodied decision-making.
Motivation
In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial understanding across heterogeneous views remains under-evaluated.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.