Benchmark Radar
AI BENCHMARK PROFILE

MindEdit-Bench

Consumer & ProductivityMultimodal PerceptionZODAOfficial

MindEdit-Bench evaluates counterfactual spatial reasoning in VLMs with 1,003 multiple-choice questions from private indoor scenes, covering six task types.

Released
2026-07-01
Readiness
Inspectable
Primary field
Consumer & Productivity

Why it matters

Tests whether VLMs can reason about hypothetical object manipulations, a capability not covered by existing benchmarks, with human-verified answers.

Motivation

Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.