Benchmark Radar
AI BENCHMARK PROFILE

PriVE-Bench

General AIMultimodal Perception

PriVE-Bench evaluates vision-language models' visual grounding using paired original and counterfactual images, with PriVE-Tools extending to tool-derived evidence.

Released
2026-07-14
Readiness
Paper only
Primary field
General AI

Why it matters

Assesses whether vision-language models rely on learned priors rather than image content, and whether additional visual tools can improve grounding.

Motivation

Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.