AI BENCHMARK PROFILE
SmellBench
SmellBench is a code refactoring benchmark that proactively injects code smells into clean code snippets. It contains 294 cases across 7 smell types, 3 difficulty levels, and 2 instruction settings, with evaluation covering functional correctness, localization, and refactoring quality.
- Released
- 2026-06-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks focus on functional correctness, not long-term maintainability. SmellBench could evaluate code agents' ability to produce maintainable code.
Motivation
Code Agents have achieved remarkable advances in recent years, exhibiting strong capabilities across a wide range of software engineering tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.