Benchmark Radar
AI BENCHMARK PROFILE

Distract-Bench

General AISafety & TrustworthinessYizheng-Sun

Distract-Bench evaluates robustness of vision-language models to semantic visual distractions, which are meaningful but task-irrelevant cues that preserve the ground-truth answer.

Released
2026-06-08
Readiness
Runnable
Primary field
General AI

Why it matters

It exposes a distinct failure mode where models perceive evidence correctly but reason from distracting cues, shifting robustness evaluation from perceptual degradation to distraction handling.

Motivation

Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable real-world application requires handling visual inputs that are messier than clean, curated benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.