SWE-Mutation
SWE-Mutation evaluates LLM-generated test suites in software engineering by using 2,636 mutated variants derived from 800 original instances across nine programming languages, measuring verification and detection rates.
- Released
- 2026-05-21
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
High-quality test suites are critical for program repair and reinforcement learning signals. SWE-Mutation reveals inadequacies in LLM-generated tests, guiding improvements in code generation and validation.
Motivation
Evaluating software engineering capabilities has become a core component of modern large language models (LLMs); however, the key bottleneck hindering further scaling lies not in the scarcity of high-quality solutions, but in the lack of high-quality test suites.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.