Benchmark Radar
AI BENCHMARK PROFILE

EHR-Complex

Health & Life SciencesAgents

EHR-Complex evaluates clinical agent performance on interactive reasoning over MIMIC-IV electronic health records. It consists of about 52K tasks across six clinical intents, requiring agents to execute SQL or Python in a sandboxed environment to answer patient- and population-level queries. Scoring is based on exact-match accuracy against expected outcomes.

Released
2026-06-22
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Existing clinical benchmarks often rely on simplified, static SQL generation, failing to reflect real-world EHR complexity. EHR-Complex introduces interactive, multi-step reasoning tasks with compositional queries, revealing that state-of-the-art models achieve only 62.3% accuracy and exhibit fragility under repeated sampling, highlighting significant room for improvement in robust clinical reasoning.

Motivation

Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often operating on idealized, clean EHRs via static SQL generation rather than interactive execution.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.