Benchmark Radar
AI BENCHMARK PROFILE

LongRCA Bench

General AIKnowledge & ReasoningLongRCA Bench Team

LongRCA Bench is a benchmark for diagnosing responsible roles and root causes in long-horizon agent failures. It comprises 1,140 failed trajectories across five domains, with human labels for the responsible role and the earliest decisive root-cause step. Evaluation focuses on responsible-role accuracy and exact root-step accuracy.

Released
2026-08-15
Readiness
Paper only
Primary field
General AI

Why it matters

Long-horizon agent failures are difficult to debug, and outcome-level metrics obscure where errors occur. LongRCA Bench provides a standardized testbed for failure attribution, enabling comparison of methods that localize root causes and assign responsibility, which is crucial for improving agent reliability.

Motivation

When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.