AI BENCHMARK PROFILE
AtmosCoder-Bench
AtmosCoder-Bench is an execution-grounded benchmark for LLMs on atmospheric science computation, with 436 problems and 3,910 variants, grading by executing code solutions against ground truth.
- Released
- 2026-08-19
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing evaluations score final answers, overlooking calculation process. AtmosCoder-Bench makes process visible, revealing failures in applying formulas and adapting to regimes.
Motivation
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.