Benchmark Radar
AI BENCHMARK PROFILE

TelemetrySuffBench

General AIMultimodal PerceptionAnonymous

TelemetrySuffBench evaluates models on fault-origin diagnosis in agent telemetry, using controlled multi-component traces with delayed-binding faults, paired coarse views, seven-factor telemetry masks, and exact-equal ambiguous origin pairs.

Released
2026-08-08
Readiness
Inspectable
Primary field
General AI

Why it matters

It addresses the evaluation gap in diagnosing failure origins from agent telemetry, showing that full telemetry yields high localization accuracy but coarse views preserve detection while limiting localization, highlighting the need for explicit decision-to-provenance links and abstention safeguards.

Motivation

Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.