AI BENCHMARK PROFILE
TS-Skill
A controlled benchmark for evaluating analytical skills in time-series question answering: temporal scale selection, temporal localization, and cross-interval integration.
- Released
- 2026-05-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Skill-level evaluation reveals temporal reasoning failures obscured by aggregate scores.
Motivation
Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.