Benchmark Radar
AI BENCHMARK PROFILE

Evo-Bench

General AIKnowledge & ReasoningRUCAIBox

Evo-Bench is a benchmark for evaluating whether language models can autonomously improve their agent harness, with 608 harness-sensitive tasks across Search, Office, and General domains, fixed policy model, and resource budget.

Released
2026-08-10
Readiness
Runnable
Primary field
General AI

Why it matters

It isolates harness evolution from base model strength, addressing a gap in evaluating agent self-improvement capabilities.

Motivation

Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.