SWE-Marathon
SWE-Marathon evaluates AI agents on 20 ultra-long-horizon software engineering tasks, each with a unique executable environment, a human-written reference solution, and a multi-layer verification suite. Tasks average 27.2M tokens per logged agent attempt, requiring sustained progress over hours and millions of tokens.
- Released
- 2026-06-05
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing agent benchmarks focus on short tasks, limiting measurement of planning, long-context understanding, and memory. SWE-Marathon addresses the gap by providing a longer-horizon evaluation that exposes practical limitations in agent autonomy and highlights failure modes like reward hacking, informing development of more robust agents.
Motivation
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.