AI BENCHMARK PROFILE
Who&When Pro
Who&When Pro is a benchmark for automated failure attribution in agentic systems, containing 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks. It uses a controlled pipeline that injects failures after replaying successful prefixes.
- Released
- 2026-07-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
As agents become more capable, automated failure attribution is crucial for debugging and safety. This benchmark provides a large-scale evaluation to guide the development of systems that can identify where and why failures occur.
Motivation
Automated failure attribution uses LLMs to identify where and why agentic systems fail.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.