SWE-Together
SWE-Together evaluates coding agents in multi-turn interactive user sessions reconstructed from real user-agent interactions. It comprises 109 repository-level tasks with a reactive LLM-based user simulator, measuring final repository correctness and the number of corrective feedback turns.
- Released
- 2026-06-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing coding-agent benchmarks often evaluate static, single-turn tasks, missing the interactive nature of real coding assistance. SWE-Together provides a reproducible protocol for assessing agents as collaborators, capturing both task success and user effort, offering practical value for comparing agents in realistic settings.
Motivation
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.