AI BENCHMARK PROFILE
IMUG-Bench
IMUG-Bench evaluates unified multimodal models on multi-turn interleaved image-text understanding and generation. It includes 3,113 samples and 12,034 interaction turns across Static Spatial, Temporal Causal, and Hybrid classes, with dynamic understanding questions.
- Released
- 2026-06-08
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks fail to evaluate multi-turn interleaved interactions and expose bias. IMUG-Bench provides a comprehensive evaluation revealing capability boundaries and failure modes, and explores test-time scaling strategies.
Motivation
In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.