Benchmark Radar
AI BENCHMARK PROFILE

AirGroundBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceAirGroundBench Team

AirGroundBench evaluates multi-view spatial intelligence in multimodal large language models through UAV-UGV collaborative tasks, including 62,000 dual-view multiple-choice questions and 115 navigation episodes across 11 simulated environments.

Released
2026-06-26
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

This benchmark addresses the gap in assessing geometric consistency across heterogeneous views, providing a structured evaluation for capabilities like cross-view alignment and spatial reasoning that are critical for embodied decision-making.

Motivation

In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial understanding across heterogeneous views remains under-evaluated.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.