AeroCopilotBench
AeroCopilotBench is a two-tier benchmark for evaluating LLM agents as aviation copilots. Tier-1 uses 1,200 multiple-choice questions for knowledge assessment, while Tier-2 includes 73 procedural tasks in an interactive virtual cockpit environment, with safety-gated evaluation.
- Released
- 2026-08-17
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Aviation evaluations often focus on static knowledge and cannot test procedural execution and safety compliance. AeroCopilotBench provides a reproducible interactive environment and safety-gated scoring to assess agents' task completion and trajectory safety.
Motivation
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.