Benchmark Radar
AI BENCHMARK PROFILE

RealBench

General AIKnowledge & ReasoningNWP-Benchmark contributors

Evaluates data-driven numerical weather forecasting models under operational conditions, using strictly out-of-distribution test data from 2025 and integrating low-latency operational analysis and large-scale in-situ observations from over 10,000 stations. It provides metrics for global forecasting and for extreme events such as heatwaves, cold surges, and tropical cyclones.

Released
2026-05-24
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks rely on reanalysis products that do not reflect real-time operational constraints, leading to mismatches between benchmark scores and real-world performance. This benchmark provides a more faithful and operationally relevant evaluation paradigm.

Motivation

Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.