AI BENCHMARK PROFILE
PriVE-Bench
PriVE-Bench evaluates vision-language models' visual grounding using paired original and counterfactual images, with PriVE-Tools extending to tool-derived evidence.
- Released
- 2026-07-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Assesses whether vision-language models rely on learned priors rather than image content, and whether additional visual tools can improve grounding.
Motivation
Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.