Active-SWE
Active-SWE evaluates coding agents on proactive bug fixing: detecting and fixing multiple bugs without issue reports. It includes 1,663 tasks across six bug categories and eight languages, with stages for recorded bugs, potential bugs, and judge validation.
- Released
- 2026-08-05
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing SWE benchmarks assume detailed issue reports are available, which is unrealistic. Active-SWE fills the gap by testing agents' ability to discover and fix bugs proactively, a capability that current state-of-the-art agents struggle with, providing a more realistic evaluation of coding agents.
Motivation
Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.