Benchmark Radar
AI BENCHMARK PROFILE

SmellBench

General AICoding & Software Engineering

SmellBench is a code refactoring benchmark that proactively injects code smells into clean code snippets. It contains 294 cases across 7 smell types, 3 difficulty levels, and 2 instruction settings, with evaluation covering functional correctness, localization, and refactoring quality.

Released
2026-06-04
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks focus on functional correctness, not long-term maintainability. SmellBench could evaluate code agents' ability to produce maintainable code.

Motivation

Code Agents have achieved remarkable advances in recent years, exhibiting strong capabilities across a wide range of software engineering tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.