Benchmark Radar
AI BENCHMARK PROFILE

TS-Skill

General AIKnowledge & Reasoning

A controlled benchmark for evaluating analytical skills in time-series question answering: temporal scale selection, temporal localization, and cross-interval integration.

Released
2026-05-23
Readiness
Paper only
Primary field
General AI

Why it matters

Skill-level evaluation reveals temporal reasoning failures obscured by aggregate scores.

Motivation

Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.