Benchmark Radar
AI BENCHMARK PROFILE

MM-CreativityBench

General AIAgentsTool CallingMM-CreativityBench team

MM-CreativityBench evaluates large multimodal models on affordance-grounded creative tool use, requiring scene inspection, entity/part selection, and physically feasible solutions in visually rich environments.

Released
2026-05-25
Readiness
Runnable
Primary field
General AI

Why it matters

It probes beyond pattern recognition, assessing grounded exploration and reasoning. Current models show gaps in exploration and hallucination, motivating preference-based alignment.

Motivation

Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded solutions in open-ended environments, beyond pattern recognition.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.