AI BENCHMARK PROFILE
MissionBench
A benchmark for mission-level evaluation of MLLMs in aerial 3D environments, consisting of 120 missions across five simulated environments and four task families.
- Released
- 2026-07-24
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Highlights the difficulty of zero-shot long-horizon embodied tasks and the need for closed-loop evaluation, motivating scaling-driven improvements.
Motivation
Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.