Benchmark Radar
AI BENCHMARK PROFILE

VisualFLIP

General AIMultimodal PerceptionVisualFLIP team

VisualFLIP evaluates multimodal LLMs on visual reasoning with 1,374 images in paired perturbation tasks. Each pair has a fixed question but minimally changed visual evidence so the answer flips. Scoring uses pair accuracy and Collapse Rate to test evidence dependence in capabilities like cardinality, attribute, spatial, and logic.

Released
2026-06-05
Readiness
Inspectable
Primary field
General AI

Why it matters

Accuracy alone can hide flawed reasoning. VisualFLIP exposes whether models truly rely on task-critical visual changes, distinguishing robust grounding from guesswork. This helps select models for high-stakes visual tasks.

Motivation

When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.