Benchmark Radar
AI BENCHMARK PROFILE

TS-Fault

General AISafety & TrustworthinessHKUST(GZ)

Evaluates time series forecasting models under four explicit structural fault modes (time-warped shock, dependency-fracture shock, regime-transition missingness, cascading sensor-to-system failure) injected into lookback windows, with paired clean/corrupt protocol and five difficulty levels across nine datasets and six domains.

Released
2026-06-16
Readiness
Runnable
Primary field
General AI

Why it matters

Standard clean-data leaderboards assume a single error metric predicts deployed reliability, but real faults are structured events. TS-Fault provides a diagnostic protocol that isolates robustness to named fault mechanisms at tunable severities, revealing that clean accuracy anti-correlates with robustness and mechanism-level faults reorder model rankings.

Motivation

Time series forecasting (TSF) underpins consequential decisions in energy, transportation, finance, and healthcare, yet TSF models are almost universally ranked by a single number (e.g., average error) on clean held-out data, under the implicit assumption that it predicts deployed reliability.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.