AI BENCHMARK PROFILE
ClawTrack
ClawTrack is a dual-assessment benchmark for agents, measuring task outcomes and process quality across 320 tasks in 8 domains with 25+ mock services.
- Released
- 2026-07-30
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It aims to decompose agent success into reasoning dimensions, which could improve attribution and post-training filtering.
Motivation
As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.