Benchmark Radar
AI BENCHMARK PROFILE

CatchBench

General AIKnowledge & ReasoningCatchBench maintainers

Assesses agent failure auditing across PRE, LIVE, and POST information states with seven task contracts covering evidential and Gold-derived diagnostics.

Released
2026-08-24
Readiness
Runnable
Primary field
General AI

Why it matters

Unifies previously separate auditing settings under one interface and exposes label-process shortcuts, helping users judge whether benchmark scores reflect reasoning.

Motivation

When can an agent failure be caught?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.