Benchmark Radar
AI BENCHMARK PROFILE

StaminaBench

General AICoding & Software Engineering

StaminaBench stress-tests coding agents over 100 interaction turns of change requests, measuring how many consecutive turns they handle before failing, with programmatically generated tasks and black-box HTTP evaluation.

Released
2026-06-17
Readiness
Paper only
Primary field
General AI

Why it matters

Fills the gap in evaluating multi-turn coding agent stamina, which is critical for real-world vibe-coding sessions that often extend over many turns.

Motivation

We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.