Benchmark Radar
AI BENCHMARK PROFILE

EHRBench

Health & Life SciencesKnowledge & Reasoning

EHRBench evaluates LLM-based clinical decision-making using nearly 1M QA items generated from real EHR trajectories, covering diagnosis, treatment, and prognosis tasks.

Released
2026-05-28
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Existing clinical decision benchmarks often lack scale and reliability; EHRBench provides a large-scale, EHR-grounded evaluation that tests models on practical inference tasks, offering insights into model capabilities and gaps for clinical deployment.

Motivation

Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future health outcomes under incomplete evidence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.