Benchmark Radar
AI BENCHMARK PROFILE

AeroCopilotBench

General AIAgents

AeroCopilotBench is a two-tier benchmark for evaluating LLM agents as aviation copilots. Tier-1 uses 1,200 multiple-choice questions for knowledge assessment, while Tier-2 includes 73 procedural tasks in an interactive virtual cockpit environment, with safety-gated evaluation.

Released
2026-08-17
Readiness
Paper only
Primary field
General AI

Why it matters

Aviation evaluations often focus on static knowledge and cannot test procedural execution and safety compliance. AeroCopilotBench provides a reproducible interactive environment and safety-gated scoring to assess agents' task completion and trajectory safety.

Motivation

Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.