Benchmark Radar
AI BENCHMARK PROFILE

MC-CXR

Health & Life SciencesMultimodal Perception

MC-CXR is a benchmark of 240 chest X-ray cases (2,522 instances) for evaluating context-induced disruption in vision-language models. It pairs reliable and misleading context across text and prior imaging, with visual overlays. Defines three task families and two paired metrics: switch-to-wrong rate and context-aligned error rate.

Released
2026-08-25
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

VLMs in clinical pipelines may be disrupted by plausible but misleading context, reducing diagnostic accuracy. MC-CXR isolates this disruption with paired perturbation, revealing high switch rates and text-visual asymmetry, supporting safer deployment.

Motivation

Vision-language models (VLMs) are increasingly used in clinical pipelines where a chest X-ray is interpreted alongside retrieved reports, preliminary notes, or prior imaging.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.