Benchmark Radar
AI BENCHMARK PROFILE

ORCA-bench

General AICoding & Software EngineeringHarbor Framework

ORCA-bench evaluates language model agents on root cause analysis in a production-fidelity oncall setting, using a live OpenTelemetry-instrumented microservice system with 1,079 tasks varying in report specificity, time-to-detection, and fault co-occurrence. Agents access metrics, logs, traces, and source code via real telemetry interfaces.

Released
2026-07-30
Readiness
Inspectable
Primary field
General AI

Why it matters

Standard coding benchmarks do not capture the complexity of oncall RCA, where agents must reason over noisy, heterogeneous data. ORCA-bench provides a reproducible testbed to measure agent readiness for production reliability tasks, revealing a significant performance gap even for frontier models.

Motivation

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.