TelemetrySuffBench
TelemetrySuffBench evaluates models on fault-origin diagnosis in agent telemetry, using controlled multi-component traces with delayed-binding faults, paired coarse views, seven-factor telemetry masks, and exact-equal ambiguous origin pairs.
- Released
- 2026-08-08
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
It addresses the evaluation gap in diagnosing failure origins from agent telemetry, showing that full telemetry yields high localization accuracy but coarse views preserve detection while limiting localization, highlighting the need for explicit decision-to-provenance links and abstention safeguards.
Motivation
Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.