Benchmark Radar
AI BENCHMARK PROFILE

SWE-Marathon

General AICoding & Software Engineeringswe-marathon.org

SWE-Marathon evaluates AI agents on 20 ultra-long-horizon software engineering tasks, each with a unique executable environment, a human-written reference solution, and a multi-layer verification suite. Tasks average 27.2M tokens per logged agent attempt, requiring sustained progress over hours and millions of tokens.

Released
2026-06-05
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing agent benchmarks focus on short tasks, limiting measurement of planning, long-context understanding, and memory. SWE-Marathon addresses the gap by providing a longer-horizon evaluation that exposes practical limitations in agent autonomy and highlights failure modes like reward hacking, informing development of more robust agents.

Motivation

AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.