EnterpriseClawBench
Evaluates coding agents on realistic enterprise workflows using 852 reproducible tasks with recovered fixtures, prompts, role classes, skill subclasses, hard rules, and semantic rubrics. Scoring covers artifact delivery, visual quality, cost, runtime, and skill transfer.
- Released
- 2026-06-22
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Enterprise agent evaluation often collapses performance into a single score, ignoring cost, runtime, and artifact quality. This benchmark provides a multi-faceted scoring contract and a reusable construction/evaluation protocol, enabling practical comparisons of harness-model systems in real workplace settings.
Motivation
Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.