Benchmark Radar
AI BENCHMARK PROFILE

ClinicalMC

Health & Life SciencesKnowledge & Reasoning

ClinicalMC evaluates LLM clinical decision-making across multi-course patient trajectories, with 1,275 Chinese and 5,804 English samples spanning four stages from admission to discharge. Includes triage, examination, diagnosis, treatment, and final diagnosis.

Released
2026-06-02
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Existing clinical benchmarks focus on single-course scenarios, missing the complexity of evolving patient conditions. ClinicalMC enables evaluation of dynamic multi-turn decision-making, supporting safer LLM deployment in healthcare.

Motivation

Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making scenarios.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.