LongRCA Bench
LongRCA Bench is a benchmark for diagnosing responsible roles and root causes in long-horizon agent failures. It comprises 1,140 failed trajectories across five domains, with human labels for the responsible role and the earliest decisive root-cause step. Evaluation focuses on responsible-role accuracy and exact root-step accuracy.
- Released
- 2026-08-15
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Long-horizon agent failures are difficult to debug, and outcome-level metrics obscure where errors occur. LongRCA Bench provides a standardized testbed for failure attribution, enabling comparison of methods that localize root causes and assign responsibility, which is crucial for improving agent reliability.
Motivation
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.