DeepSWE
Evaluates coding agents on 113 original, long-horizon software engineering tasks across 91 open-source repositories in five languages. Tasks are written from scratch, with hand-written verifiers that check requested functionality and accept any correct implementation.
- Released
- 2026-07-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the gap of benchmarks relying on mined fixes and inherited tests, which can overstate model capability due to pretraining exposure and rigid grading. Provides a reusable evaluation path with verifiers and trajectories for assessing genuine problem-solving ability.
Motivation
DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.