Benchmark Radar
AI BENCHMARK PROFILE

MMOOC

General AIMultimodal PerceptionMMOOC Project Team

MMOOC evaluates multimodal large language models on out-of-context (OOC) and shifted in-context (Shifted IC) visual question answering. It contains over 41K image-question pairs covering three question formats, eight shift types, and six visual scenarios. Responses are scored for accuracy and refusal rate, with an LLM-as-a-judge metric for reasoning correctness.

Released
2026-07-30
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on unanswerable questions but overlook answerable shifted contexts. MMOOC provides a joint evaluation of refusal and robust answering, offering insight into model reliability in real-world scenarios where contexts are imperfect. This supports comparing models on a balanced measure of capability and safety.

Motivation

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.