Benchmark Radar
AI BENCHMARK PROFILE

DeepSWE

General AICoding & Software EngineeringDataCurve

Evaluates coding agents on 113 original, long-horizon software engineering tasks across 91 open-source repositories in five languages. Tasks are written from scratch, with hand-written verifiers that check requested functionality and accept any correct implementation.

Released
2026-07-08
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap of benchmarks relying on mined fixes and inherited tests, which can overstate model capability due to pretraining exposure and rigid grading. Provides a reusable evaluation path with verifiers and trajectories for assessing genuine problem-solving ability.

Motivation

DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.