Benchmark Radar
AI BENCHMARK PROFILE

VES-Bench

General AIMultimodal PerceptionSearch & Retrieval

Tests long-horizon video understanding with 600 Temporal Ordering and Event Counting questions, auditing whether decoded frames cover all evidence intervals.

Released
2026-08-23
Readiness
Paper only
Primary field
General AI

Why it matters

Shifts evaluation from final answers to evidence coverage, revealing whether correct answers rest on complete observation of long videos.

Motivation

A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.