Benchmark Radar
AI BENCHMARK PROFILE

ClinHallu

Health & Life SciencesMultimodal PerceptionAlibaba DAMO Academy

ClinHallu is a benchmark for diagnosing stage-wise hallucinations in medical multimodal large language models. It contains 7,031 instances with structured reasoning traces decomposed into visual recognition, knowledge recall, and reasoning integration, along with stage-replacement interventions for measuring final answer changes.

Released
2026-06-12
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Existing medical hallucination benchmarks often ignore the source of hallucinations within reasoning. ClinHallu allows for fine-grained diagnosis of where errors originate, providing a testbed for improving model reliability in clinical decision support.

Motivation

Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.