AI BENCHMARK PROFILE
Distract-Bench
Distract-Bench evaluates robustness of vision-language models to semantic visual distractions, which are meaningful but task-irrelevant cues that preserve the ground-truth answer.
- Released
- 2026-06-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It exposes a distinct failure mode where models perceive evidence correctly but reason from distracting cues, shifting robustness evaluation from perceptual degradation to distraction handling.
Motivation
Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable real-world application requires handling visual inputs that are messier than clean, curated benchmarks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.