Benchmark Radar
AI BENCHMARK PROFILE

SWE-NFI

General AICoding & Software Engineering

A benchmark of 188 tasks for evaluating coding agents on non-functional improvements in Python projects, with 92 executable rules combining functional correctness and rule-based evaluation.

Released
2026-07-29
Readiness
Paper only
Primary field
General AI

Why it matters

Existing coding benchmarks focus on functional correctness; this benchmark addresses the gap in evaluating behavior-preserving code quality improvements, useful for assessing real-world software engineering capabilities.

Motivation

Although coding agents have achieved impressive performance on correctness-oriented benchmarks, their ability to make behavior-preserving non-functional improvements (NFIs) remains underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.