Benchmark Radar
AI BENCHMARK PROFILE

TimeSage-MT

General AIKnowledge & ReasoningTimeSage-MT Team

TimeSage-MT evaluates agentic time series reasoning in multi-turn dialogues, covering 240 tasks and 2,680 turns across 8 domains, with a protocol and leaderboard for comparing systems.

Released
2026-05-31
Readiness
Paper only
Primary field
General AI

Why it matters

It fills a gap in benchmarking agentic time series analysis, which differs from single-step tasks, offering a way to assess multi-turn memory, uncertainty handling, and decision-making, with practical value for developing and comparing LLM agents.

Motivation

Time series data inform critical decisions across many real-world domains.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.