SHOVIR
SHOVIR evaluates vision shortcut learning in radiology report generation by extending MIMIC-CXR and PadChest-GR with per-box CheXpert labels. It defines image-level and disease-level occlusion experiments that compare model predictions on clean images against localized perturbations to isolate direct and contextual shortcut failures at the disease-class level.
- Released
- 2026-06-29
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Standard RRG metrics miss whether diagnostic statements are grounded in actual image evidence, allowing models to exploit dataset priors. SHOVIR provides a protocol to assess spatial grounding, revealing that high report quality can coexist with shallow visual reliance, which is critical for clinical deployment decisions.
Motivation
Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.