Benchmark Radar
AI BENCHMARK PROFILE

ClawProBench

General AIKnowledge & Reasoning

Evaluates agent configurations on 102 live-runtime and 68 frozen holdout scenarios, scoring execution traces with a safety-gated formula covering correctness, process quality, and efficiency.

Released
2026-08-23
Readiness
Paper only
Primary field
General AI

Why it matters

It moves agent evaluation beyond final answers to expose runtime-navigation, safety-boundary, and repeated-execution failures that conventional leaderboards hide.

Motivation

Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.