AI BENCHMARK PROFILE
Video-MME-Logical
Evaluates video temporal-logical reasoning in multimodal LLMs across 25 fine-grained task categories, with difficulty-controlled final-answer scoring and intermediate-state diagnostics.
- Released
- 2026-06-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Isolates temporal-logical capabilities from static recognition, revealing significant human-model gaps and providing a scalable testbed for analysis.
Motivation
Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events in individual frames?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.