Benchmark Radar
AI BENCHMARK PROFILE

IMUG-Bench

General AIMultimodal Perception

IMUG-Bench evaluates unified multimodal models on multi-turn interleaved image-text understanding and generation. It includes 3,113 samples and 12,034 interaction turns across Static Spatial, Temporal Causal, and Hybrid classes, with dynamic understanding questions.

Released
2026-06-08
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks fail to evaluate multi-turn interleaved interactions and expose bias. IMUG-Bench provides a comprehensive evaluation revealing capability boundaries and failure modes, and explores test-time scaling strategies.

Motivation

In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.