MMGist
MMGist is a curated multimodal benchmark covering seven capability dimensions with 7,262 items, derived from 18 existing benchmarks via filtering pipelines. It evaluates vision-language models across dimensions such as Visual Logic and Expert Knowledge.
- Released
- 2026-06-21
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
MMGist addresses the need for reliable and discriminative evaluation of vision-language models, aiming to reduce redundancy and saturation in existing benchmarks while preserving model rankings.
Motivation
We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effectively measure multimodal understanding; 2) many items are already close to performance saturation for current LVLMs, which limits their discriminative power; 3) a small number of anomalous items affect the reliability of evaluation results.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.