AI BENCHMARK PROFILE
ClawProBench
Evaluates agent configurations on 102 live-runtime and 68 frozen holdout scenarios, scoring execution traces with a safety-gated formula covering correctness, process quality, and efficiency.
- Released
- 2026-08-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It moves agent evaluation beyond final answers to expose runtime-navigation, safety-boundary, and repeated-execution failures that conventional leaderboards hide.
Motivation
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.