TimeSage-EV
TimeSage-EV evaluates LLM agents on time series analysis tasks in evolving environments, using 60 institutional scenarios across 6 domains with 1,485 scenario-period QA pairs. Agents receive data and reports, with withheld target releases as ground truth, assessing state identification, data summarization, and outlook reasoning.
- Released
- 2026-08-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing time series QA benchmarks rely on fixed snapshots, but real-world data is released periodically, affecting conclusions. TimeSage-EV fills this gap by evaluating agentic temporal validity and cutoff-aware evidence use, providing a decision-value for deploying LLM agents in high-stakes domains where data updates matter.
Motivation
Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.