Benchmark Radar
AI BENCHMARK PROFILE

Who&When Pro

General AIKnowledge & Reasoning

Who&When Pro is a benchmark for automated failure attribution in agentic systems, containing 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks. It uses a controlled pipeline that injects failures after replaying successful prefixes.

Released
2026-07-10
Readiness
Paper only
Primary field
General AI

Why it matters

As agents become more capable, automated failure attribution is crucial for debugging and safety. This benchmark provides a large-scale evaluation to guide the development of systems that can identify where and why failures occur.

Motivation

Automated failure attribution uses LLMs to identify where and why agentic systems fail.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.