AI BENCHMARK PROFILE
IRTS-ToolBench
IRTS-ToolBench is a benchmark of 1,700 questions across 10 task types and 13 domains for evaluating irregular univariate time-series question answering, with standardized inputs and a reproducible evaluation protocol.
- Released
- 2026-06-13
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing TSQA benchmarks assume regular sampling, leaving a gap for real-world irregular data. This benchmark provides standardized evaluation for LLMs and AI agents on irregular time series, with golden tool sets for tool-selection analysis.
Motivation
Time series data in real-world deployments is overwhelmingly irregular.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.