AI BENCHMARK PROFILE
FaulT-Bench
A benchmark of 200 troubleshooting scenarios across eight network topologies evaluating network troubleshooting LLM agents under unreliable user tickets, including false fault reports and incorrect device attribution.
- Released
- 2026-08-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Evaluates the robustness of network troubleshooting agents against noisy, real-world tickets rather than only accurate inputs.
Motivation
LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.