AI BENCHMARK PROFILE
AGENTCHAOSBENCH
AGENTCHAOSBENCH is a dataset of sanitized execution traces from five agentic applications with injected runtime faults, used to evaluate fault detection and localization from telemetry.
- Released
- 2026-08-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses runtime fault diagnosis in LLM agentic systems, offering a reproducible task for comparing diagnostic methods across tool and agent boundaries.
Motivation
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.