SWE-INTERACT
SWE-Interact evaluates coding agents on multi-turn, interactive software engineering tasks where a simulated user provides vague instructions, reveals requirements progressively, and gives feedback. The benchmark comprises 75 tasks and measures agents' ability to discover user intent, adapt to evolving requirements, and build on prior work.
- Released
- 2026-06-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing SWE benchmarks focus on single-turn autonomous implementation, but real developer workflows are interactive. SWE-Interact fills the gap by measuring performance on long-horizon, user-driven tasks, showing that strong single-turn performance does not reliably transfer. This provides a more realistic evaluation axis for coding agents and guides development of models that can collaborate effectively with users.
Motivation
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.