Benchmark Radar
AI BENCHMARK PROFILE

AtmosCoder-Bench

General AIKnowledge & Reasoningacodercat

AtmosCoder-Bench is an execution-grounded benchmark for LLMs on atmospheric science computation, with 436 problems and 3,910 variants, grading by executing code solutions against ground truth.

Released
2026-08-19
Readiness
Runnable
Primary field
General AI

Why it matters

Existing evaluations score final answers, overlooking calculation process. AtmosCoder-Bench makes process visible, revealing failures in applying formulas and adapting to regimes.

Motivation

Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.