Benchmark Radar
AI BENCHMARK PROFILE

MissionBench

Robotics & Autonomous SystemsMultimodal Perception

A benchmark for mission-level evaluation of MLLMs in aerial 3D environments, consisting of 120 missions across five simulated environments and four task families.

Released
2026-07-24
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Highlights the difficulty of zero-shot long-horizon embodied tasks and the need for closed-loop evaluation, motivating scaling-driven improvements.

Motivation

Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.