AI BENCHMARK PROFILE
DiagChain
DiagChain evaluates LLM agents on evidence-grounded attack chain reconstruction from heterogeneous telemetry. It provides 69 scenarios and five metrics for stage-wise assessment of reasoning steps.
- Released
- 2026-08-04
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Existing benchmarks often report end-to-end accuracy, hiding where errors arise. DiagChain enables systematic diagnosis of intermediate stages, offering actionable insights for improving cybersecurity agents.
Motivation
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.