Benchmark Radar
AI BENCHMARK PROFILE

The Image Reconstruction Game

General AIMultimodal Perception

A benchmark for iterative multimodal dialogue where a vision-language model issues corrective instructions to an image generator over multiple turns, with accumulated common ground observable as a rendered image.

Released
2026-06-01
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the evaluation of interactive language-vision systems and common ground building, but lacks a clear public comparison path.

Motivation

We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image generator across multiple turns, making accumulated common ground directly observable as a rendered image.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.