AI BENCHMARK PROFILE
MM-CreativityBench
MM-CreativityBench evaluates large multimodal models on affordance-grounded creative tool use, requiring scene inspection, entity/part selection, and physically feasible solutions in visually rich environments.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It probes beyond pattern recognition, assessing grounded exploration and reasoning. Current models show gaps in exploration and hallucination, motivating preference-based alignment.
Motivation
Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded solutions in open-ended environments, beyond pattern recognition.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.