Benchmark Radar
AI BENCHMARK PROFILE

SWE-Bench ProMax

General AICoding & Software EngineeringSWE-Bench-ProMax Team

A multilingual code refactoring benchmark with 170 instances drawn from real commits across seven programming languages. Evaluates AI agents on large-scale refactoring tasks averaging 11.4 modified files and 261.6 lines of code, using manually curated issue descriptions and test suites.

Released
2026-08-10
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing software engineering benchmarks face saturation and quality issues, with flawed tests and training data leakage. This benchmark provides a more challenging and realistic refactoring task set with rigorous curation, offering a robust measure of agent capability for long-horizon coding tasks.

Motivation

As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and th…

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.