Benchmark Radar
AI BENCHMARK PROFILE

EvoClawBench

General AIKnowledge & Reasoning

EvoClawBench evaluates whether an agent runtime can convert evidence from its own runs into reusable skills that improve fresh executions. It covers 100 tasks and 502 sub-problems across coding, data, office, security, operations, and domain-document workflows, comparing direct execution, pre-authored skills, and post-run skill summarization.

Released
2026-06-23
Readiness
Paper only
Primary field
General AI

Why it matters

The evaluation gap is assessing closed-loop skill learning in agents, where benefits are selective and cost-sensitive rather than automatic. The benchmark provides a decision value for runtime developers and users considering skill-authoring loops.

Motivation

Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skills that improve fresh executions after authoring overhead.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.