ClinHallu
ClinHallu is a benchmark for diagnosing stage-wise hallucinations in medical multimodal large language models. It contains 7,031 instances with structured reasoning traces decomposed into visual recognition, knowledge recall, and reasoning integration, along with stage-replacement interventions for measuring final answer changes.
- Released
- 2026-06-12
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Existing medical hallucination benchmarks often ignore the source of hallucinations within reasoning. ClinHallu allows for fine-grained diagnosis of where errors originate, providing a testbed for improving model reliability in clinical decision support.
Motivation
Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.