VideoOdyssey
VideoOdyssey evaluates models on ultra-long-context video understanding using videos averaging 109 minutes across 11 domains, with two subsets for visual and audio-visual understanding. Tasks include question answering with continuous certificates averaging 16 and 12.8 minutes respectively.
- Released
- 2026-05-21
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing long-video benchmarks often only test short segments, failing to capture the cognitive load of continuous reasoning over long spans. VideoOdyssey's multi-level continuous certificates provide a diagnostic for evaluating model performance across varying context lengths, addressing a gap in measuring true long-context and omni-modal understanding.
Motivation
Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.