Benchmark Radar
AI BENCHMARK PROFILE

SWE-Mutation

General AICoding & Software EngineeringSWE-Mutation Team

SWE-Mutation evaluates LLM-generated test suites in software engineering by using 2,636 mutated variants derived from 800 original instances across nine programming languages, measuring verification and detection rates.

Released
2026-05-21
Readiness
Paper only
Primary field
General AI

Why it matters

High-quality test suites are critical for program repair and reinforcement learning signals. SWE-Mutation reveals inadequacies in LLM-generated tests, guiding improvements in code generation and validation.

Motivation

Evaluating software engineering capabilities has become a core component of modern large language models (LLMs); however, the key bottleneck hindering further scaling lies not in the scarcity of high-quality solutions, but in the lack of high-quality test suites.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.