Benchmark Radar
AI BENCHMARK PROFILE

SpatialGen-Bench

General AIMultimodal Perception

ProVisE is a framework for evaluating image-generation models on spatial benchmarks by converting visual answers into structured predictions. SpatialGen-Bench is a diagnostic dataset of 470 samples across 14 spatial subtasks.

Released
2026-07-23
Readiness
Runnable
Primary field
General AI

Why it matters

Existing spatial benchmarks rely on text or coordinates, limiting image-generation models. ProVisE adapts visual answers to original metrics, enabling comparison. However, unclear if it is a standalone benchmark or a framework.

Motivation

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.