AI BENCHMARK PROFILE
StaminaBench
StaminaBench stress-tests coding agents over 100 interaction turns of change requests, measuring how many consecutive turns they handle before failing, with programmatically generated tasks and black-box HTTP evaluation.
- Released
- 2026-06-17
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Fills the gap in evaluating multi-turn coding agent stamina, which is critical for real-world vibe-coding sessions that often extend over many turns.
Motivation
We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.