Pre-Flight
Pre-Flight evaluates large language models on aviation operational knowledge via 300 multiple-choice questions drawn from international standards and airport ground operations material, covering ground operations, ICAO and FAA regulations, general aviation knowledge, and operational scenarios. Scoring is by accuracy under a standard multiple-choice protocol using the Inspect framework.
- Released
- 2026-07-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
General-purpose benchmarks do not assess aviation-specific operational safety knowledge, a high-stakes domain where incorrect reasoning can have serious consequences. This benchmark provides a domain-specific evaluation to gauge model reliability for non-safety-critical aviation operations.
Motivation
Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing assistants.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.