IDEAL-Bench
IDEAL-Bench evaluates Vision-Language Models on holistic 3D layout inference from single images of indoor scenes, scoring predictions across five numerical dimensions and a perceptual render-and-compare protocol. It uses a procedurally generated dataset of 1,000 re-renderable Blender environments.
- Released
- 2026-07-03
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Current VLM spatial evaluation relies on question answering, which misses structural understanding. IDEAL-Bench provides a reproducible, quantitative assessment of geometric and structural competencies, revealing model weaknesses in measuring scenes rather than describing them, and offering a diagnostic for genuine spatial intelligence.
Motivation
Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary axis of spatial competence under-evaluated: holistic 3D layout inference, which predicts every visible object's pose and extent from a single image in a structured form.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.