Benchmark Radar
AI BENCHMARK PROFILE

FaulT-Bench

General AIKnowledge & Reasoning

A benchmark of 200 troubleshooting scenarios across eight network topologies evaluating network troubleshooting LLM agents under unreliable user tickets, including false fault reports and incorrect device attribution.

Released
2026-08-27
Readiness
Paper only
Primary field
General AI

Why it matters

Evaluates the robustness of network troubleshooting agents against noisy, real-world tickets rather than only accurate inputs.

Motivation

LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.