Benchmark Radar
AI BENCHMARK PROFILE

Pre-Flight

General AIKnowledge & Reasoning

Pre-Flight evaluates large language models on aviation operational knowledge via 300 multiple-choice questions drawn from international standards and airport ground operations material, covering ground operations, ICAO and FAA regulations, general aviation knowledge, and operational scenarios. Scoring is by accuracy under a standard multiple-choice protocol using the Inspect framework.

Released
2026-07-02
Readiness
Paper only
Primary field
General AI

Why it matters

General-purpose benchmarks do not assess aviation-specific operational safety knowledge, a high-stakes domain where incorrect reasoning can have serious consequences. This benchmark provides a domain-specific evaluation to gauge model reliability for non-safety-critical aviation operations.

Motivation

Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing assistants.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.