AI BENCHMARK PROFILE
TimeSage-MT
TimeSage-MT evaluates agentic time series reasoning in multi-turn dialogues, covering 240 tasks and 2,680 turns across 8 domains, with a protocol and leaderboard for comparing systems.
- Released
- 2026-05-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It fills a gap in benchmarking agentic time series analysis, which differs from single-step tasks, offering a way to assess multi-turn memory, uncertainty handling, and decision-making, with practical value for developing and comparing LLM agents.
Motivation
Time series data inform critical decisions across many real-world domains.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.