Benchmark Radar
AI BENCHMARK PROFILE

EHRNote-ChatQA

Health & Life SciencesKnowledge & Reasoning

EHRNote-ChatQA is a benchmark for evidence-grounded multi-turn clinical QA over longitudinal discharge summaries. Built from MIMIC-IV, it includes 967 patient-level samples and 16,072 expert-verified QA pairs across eight clinical categories, with evidence-grounding QA pairs.

Released
2026-06-14
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

EHRNote-ChatQA addresses the gap in evaluating clinical QA systems on multi-turn, evidence-grounded reasoning over multiple documents, which reflects real clinical review workflows. It provides a rigorous benchmark for assessing evidence grounding and the compounding of errors over turns.

Motivation

Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.