AI BENCHMARK PROFILE
Evo-Bench
Evo-Bench is a benchmark for evaluating whether language models can autonomously improve their agent harness, with 608 harness-sensitive tasks across Search, Office, and General domains, fixed policy model, and resource budget.
- Released
- 2026-08-10
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It isolates harness evolution from base model strength, addressing a gap in evaluating agent self-improvement capabilities.
Motivation
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.