Benchmark Radar
AI BENCHMARK PROFILE

EvoAgentBench

General AICoding & Software Engineering

EvoAgentBench evaluates agent self-evolution via ability transfer across web research, algorithmic reasoning, software engineering, and knowledge work, using trace-grounded abilities and domain graphs with train/test splits.

Released
2026-07-06
Readiness
Inspectable
Primary field
General AI

Why it matters

Current evaluations don't isolate procedural reuse. This benchmark enables fine-grained diagnosis of experience encoding, routing, and uptake in agent self-evolution.

Motivation

Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.