AI BENCHMARK PROFILE
SWE-bench Science
A repository-level benchmark with 119 tasks from 98 GitHub repositories across 20 scientific domains, organized into issue-driven, expert-exploratory, and engineering-integration paradigms.
- Released
- 2026-08-20
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Evaluates coding agents on scientific software failures, identifying failure modes and the nuanced role of scientific knowledge in repair tasks.
Motivation
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.