Benchmark Radar
AI BENCHMARK PROFILE

SWE-bench Science

General AICoding & Software Engineering

A repository-level benchmark with 119 tasks from 98 GitHub repositories across 20 scientific domains, organized into issue-driven, expert-exploratory, and engineering-integration paradigms.

Released
2026-08-20
Readiness
Paper only
Primary field
General AI

Why it matters

Evaluates coding agents on scientific software failures, identifying failure modes and the nuanced role of scientific knowledge in repair tasks.

Motivation

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.