Benchmark Radar
AI BENCHMARK PROFILE

Video-MME-Logical

General AIMultimodal PerceptionMrakas

Evaluates video temporal-logical reasoning in multimodal LLMs across 25 fine-grained task categories, with difficulty-controlled final-answer scoring and intermediate-state diagnostics.

Released
2026-06-26
Readiness
Runnable
Primary field
General AI

Why it matters

Isolates temporal-logical capabilities from static recognition, revealing significant human-model gaps and providing a scalable testbed for analysis.

Motivation

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events in individual frames?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.