Benchmark Radar
AI BENCHMARK PROFILE

SHOVIR

Health & Life SciencesMultimodal Perception

SHOVIR evaluates vision shortcut learning in radiology report generation by extending MIMIC-CXR and PadChest-GR with per-box CheXpert labels. It defines image-level and disease-level occlusion experiments that compare model predictions on clean images against localized perturbations to isolate direct and contextual shortcut failures at the disease-class level.

Released
2026-06-29
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Standard RRG metrics miss whether diagnostic statements are grounded in actual image evidence, allowing models to exploit dataset priors. SHOVIR provides a protocol to assess spatial grounding, revealing that high report quality can coexist with shallow visual reliance, which is critical for clinical deployment decisions.

Motivation

Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.