Benchmark Radar
AI BENCHMARK PROFILE

ChainSWE

General AICoding & Software Engineering

ChainSWE evaluates coding agents on sequential, dependent bug fixes within a shared codebase. It includes 304 issues across 54 Python projects, forming chronological chains, and measures performance drop as chain length increases.

Released
2026-07-01
Readiness
Paper only
Primary field
General AI

Why it matters

Real-world software maintenance involves streams of related defects, but existing benchmarks evaluate one bug at a time. ChainSWE fills this gap by benchmarking agents on continuous workflows, revealing significant performance degradation on longer chains.

Motivation

Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.