Benchmark Radar
AI BENCHMARK PROFILE

SWE Refactor Bench

General AICoding & Software EngineeringSWE Refactor Bench Team

SWE Refactor Bench is a benchmark for evaluating coding agents on whole-repository stack migrations. It comprises 20 migrations covering 4 types of technical debt, with a three-stage evaluation protocol measuring migration completeness and behavioral correctness: Migration Audit, Behavioral Tests, and Agentic Verification.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the gap in evaluating code migration beyond mere test passing, preventing shortcut solutions. It provides a rigorous testbed for developing reliable coding agents for long-horizon refactoring tasks, with findings on agent capability across migration categories.

Motivation

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.