NatureBench
NatureBench evaluates AI coding agents on 90 tasks distilled from Nature-family publications across 6 scientific domains, scoring against each paper's reported state of the art.
- Released
- 2026-06-23
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a standardized environment and public leaderboard to measure whether coding agents can achieve discovery-level performance on real scientific problems, addressing environment-fragmentation issues in prior benchmarks.
Motivation
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction toward discovery on real scientific problems.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.