Benchmark Radar
AI BENCHMARK PROFILE

VideoOdyssey

General AIMultimodal PerceptionLong Context & Memory

VideoOdyssey evaluates models on ultra-long-context video understanding using videos averaging 109 minutes across 11 domains, with two subsets for visual and audio-visual understanding. Tasks include question answering with continuous certificates averaging 16 and 12.8 minutes respectively.

Released
2026-05-21
Readiness
Paper only
Primary field
General AI

Why it matters

Existing long-video benchmarks often only test short segments, failing to capture the cognitive load of continuous reasoning over long spans. VideoOdyssey's multi-level continuous certificates provide a diagnostic for evaluating model performance across varying context lengths, addressing a gap in measuring true long-context and omni-modal understanding.

Motivation

Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.