EvoClawBench
EvoClawBench evaluates whether an agent runtime can convert evidence from its own runs into reusable skills that improve fresh executions. It covers 100 tasks and 502 sub-problems across coding, data, office, security, operations, and domain-document workflows, comparing direct execution, pre-authored skills, and post-run skill summarization.
- Released
- 2026-06-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The evaluation gap is assessing closed-loop skill learning in agents, where benefits are selective and cost-sensitive rather than automatic. The benchmark provides a decision value for runtime developers and users considering skill-authoring loops.
Motivation
Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skills that improve fresh executions after authoring overhead.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.