Benchmark Radar
AI BENCHMARK PROFILE

AGENTCHAOSBENCH

General AIKnowledge & Reasoning

AGENTCHAOSBENCH is a dataset of sanitized execution traces from five agentic applications with injected runtime faults, used to evaluate fault detection and localization from telemetry.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses runtime fault diagnosis in LLM agentic systems, offering a reproducible task for comparing diagnostic methods across tool and agent boundaries.

Motivation

Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.