Benchmark Radar
AI BENCHMARK PROFILE

TimeVista

General AIMultimodal Perception

TimeVista is a benchmark for evaluating time series forecasting using Vision-Language Models (VLMs) as judges, with 5563 samples and rubrics for micro- and macro-level judgments.

Released
2026-06-15
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses limitations of point-wise metrics in time series forecasting, offering a human-aligned evaluation approach that could guide model selection and improvement.

Motivation

High-quality time series forecasting is pivotal for real-world decision-making.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.