AI BENCHMARK PROFILE
AgentRelBench
AgentRelBench measures repeated agent runs in a fixed task suite, computing severity-weighted damage from database state diffs with no LLM judging.
- Released
- 2026-08-15
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It provides ground-truth reliability signals that evade single-run safety audits, helping evaluators detect stochastic agent failures.
Motivation
We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.