Benchmark Radar
AI BENCHMARK PROFILE

IRTS-ToolBench

General AIKnowledge & Reasoning

IRTS-ToolBench is a benchmark of 1,700 questions across 10 task types and 13 domains for evaluating irregular univariate time-series question answering, with standardized inputs and a reproducible evaluation protocol.

Released
2026-06-13
Readiness
Runnable
Primary field
General AI

Why it matters

Existing TSQA benchmarks assume regular sampling, leaving a gap for real-world irregular data. This benchmark provides standardized evaluation for LLMs and AI agents on irregular time series, with golden tool sets for tool-selection analysis.

Motivation

Time series data in real-world deployments is overwhelmingly irregular.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.