AI BENCHMARK PROFILE
EHRBench
EHRBench evaluates LLM-based clinical decision-making using nearly 1M QA items generated from real EHR trajectories, covering diagnosis, treatment, and prognosis tasks.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Existing clinical decision benchmarks often lack scale and reliability; EHRBench provides a large-scale, EHR-grounded evaluation that tests models on practical inference tasks, offering insights into model capabilities and gaps for clinical deployment.
Motivation
Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future health outcomes under incomplete evidence.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.