AI BENCHMARK PROFILE
MindEdit-Bench
MindEdit-Bench evaluates counterfactual spatial reasoning in VLMs with 1,003 multiple-choice questions from private indoor scenes, covering six task types.
- Released
- 2026-07-01
- Readiness
- Inspectable
- Primary field
- Consumer & Productivity
Why it matters
Tests whether VLMs can reason about hypothetical object manipulations, a capability not covered by existing benchmarks, with human-verified answers.
Motivation
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.