Benchmark Radar
AI BENCHMARK PROFILE

SWE-Together

General AIAgentsCoding & Software EngineeringTogetherbench

SWE-Together evaluates coding agents in multi-turn interactive user sessions reconstructed from real user-agent interactions. It comprises 109 repository-level tasks with a reactive LLM-based user simulator, measuring final repository correctness and the number of corrective feedback turns.

Released
2026-06-29
Readiness
Runnable
Primary field
General AI

Why it matters

Existing coding-agent benchmarks often evaluate static, single-turn tasks, missing the interactive nature of real coding assistance. SWE-Together provides a reproducible protocol for assessing agents as collaborators, capturing both task success and user effort, offering practical value for comparing agents in realistic settings.

Motivation

Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.