Benchmark Radar
AI BENCHMARK PROFILE

ClawTrack

General AIAgents

ClawTrack is a dual-assessment benchmark for agents, measuring task outcomes and process quality across 320 tasks in 8 domains with 25+ mock services.

Released
2026-07-30
Readiness
Paper only
Primary field
General AI

Why it matters

It aims to decompose agent success into reasoning dimensions, which could improve attribution and post-training filtering.

Motivation

As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.