Moment-Video
Moment-Video evaluates video MLLMs on momentary visual event understanding through 1,000 human-verified video-QA pairs across four task types: temporal occurrence, temporal counting, action description, and temporal reasoning.
- Released
- 2026-06-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses the gap in evaluating models' ability to capture brief answer-critical visual evidence, which is common in practical video understanding and not covered by general video benchmarks. The results show significant room for improvement, offering a diagnostic tool for model development.
Motivation
Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.