AI BENCHMARK PROFILE
AFDBench
Evaluates generative meteorological reasoning through 7,732 expert-written forecast discussions paired with AI weather inputs, using metrics for numerical accuracy, style, and grounding.
- Released
- 2026-08-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides a high-stakes domain benchmark for factual accuracy and professional style in weather text generation, crucial for reliable AI-assisted communication.
Motivation
Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.