VisualFLIP
VisualFLIP evaluates multimodal LLMs on visual reasoning with 1,374 images in paired perturbation tasks. Each pair has a fixed question but minimally changed visual evidence so the answer flips. Scoring uses pair accuracy and Collapse Rate to test evidence dependence in capabilities like cardinality, attribute, spatial, and logic.
- Released
- 2026-06-05
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Accuracy alone can hide flawed reasoning. VisualFLIP exposes whether models truly rely on task-critical visual changes, distinguishing robust grounding from guesswork. This helps select models for high-stakes visual tasks.
Motivation
When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.