AI BENCHMARK PROFILE
SWE-NFI
A benchmark of 188 tasks for evaluating coding agents on non-functional improvements in Python projects, with 92 executable rules combining functional correctness and rule-based evaluation.
- Released
- 2026-07-29
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing coding benchmarks focus on functional correctness; this benchmark addresses the gap in evaluating behavior-preserving code quality improvements, useful for assessing real-world software engineering capabilities.
Motivation
Although coding agents have achieved impressive performance on correctness-oriented benchmarks, their ability to make behavior-preserving non-functional improvements (NFIs) remains underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.