ChainSWE
ChainSWE evaluates coding agents on sequential, dependent bug fixes within a shared codebase. It includes 304 issues across 54 Python projects, forming chronological chains, and measures performance drop as chain length increases.
- Released
- 2026-07-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Real-world software maintenance involves streams of related defects, but existing benchmarks evaluate one bug at a time. ChainSWE fills this gap by benchmarking agents on continuous workflows, revealing significant performance degradation on longer chains.
Motivation
Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.