SWE-Bench ProMax
A multilingual code refactoring benchmark with 170 instances drawn from real commits across seven programming languages. Evaluates AI agents on large-scale refactoring tasks averaging 11.4 modified files and 261.6 lines of code, using manually curated issue descriptions and test suites.
- Released
- 2026-08-10
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing software engineering benchmarks face saturation and quality issues, with flawed tests and training data leakage. This benchmark provides a more challenging and realistic refactoring task set with rigorous curation, offering a robust measure of agent capability for long-horizon coding tasks.
Motivation
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and th…
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.