ORCA-bench
ORCA-bench evaluates language model agents on root cause analysis in a production-fidelity oncall setting, using a live OpenTelemetry-instrumented microservice system with 1,079 tasks varying in report specificity, time-to-detection, and fault co-occurrence. Agents access metrics, logs, traces, and source code via real telemetry interfaces.
- Released
- 2026-07-30
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Standard coding benchmarks do not capture the complexity of oncall RCA, where agents must reason over noisy, heterogeneous data. ORCA-bench provides a reproducible testbed to measure agent readiness for production reliability tasks, revealing a significant performance gap even for frontier models.
Motivation
Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.