AI BENCHMARK PROFILE
CatchBench
Assesses agent failure auditing across PRE, LIVE, and POST information states with seven task contracts covering evidential and Gold-derived diagnostics.
- Released
- 2026-08-24
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Unifies previously separate auditing settings under one interface and exposes label-process shortcuts, helping users judge whether benchmark scores reflect reasoning.
Motivation
When can an agent failure be caught?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.