MMOOC
MMOOC evaluates multimodal large language models on out-of-context (OOC) and shifted in-context (Shifted IC) visual question answering. It contains over 41K image-question pairs covering three question formats, eight shift types, and six visual scenarios. Responses are scored for accuracy and refusal rate, with an LLM-as-a-judge metric for reasoning correctness.
- Released
- 2026-07-30
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks focus on unanswerable questions but overlook answerable shifted contexts. MMOOC provides a joint evaluation of refusal and robust answering, offering insight into model reliability in real-world scenarios where contexts are imperfect. This supports comparing models on a balanced measure of capability and safety.
Motivation
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.