Benchmark Radar
AI BENCHMARK PROFILE

AgentRelBench

General AIAgentsAgentRelBench team

AgentRelBench measures repeated agent runs in a fixed task suite, computing severity-weighted damage from database state diffs with no LLM judging.

Released
2026-08-15
Readiness
Runnable
Primary field
General AI

Why it matters

It provides ground-truth reliability signals that evade single-run safety audits, helping evaluators detect stochastic agent failures.

Motivation

We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.