EHR-Complex
EHR-Complex evaluates clinical agent performance on interactive reasoning over MIMIC-IV electronic health records. It consists of about 52K tasks across six clinical intents, requiring agents to execute SQL or Python in a sandboxed environment to answer patient- and population-level queries. Scoring is based on exact-match accuracy against expected outcomes.
- Released
- 2026-06-22
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Existing clinical benchmarks often rely on simplified, static SQL generation, failing to reflect real-world EHR complexity. EHR-Complex introduces interactive, multi-step reasoning tasks with compositional queries, revealing that state-of-the-art models achieve only 62.3% accuracy and exhibit fragility under repeated sampling, highlighting significant room for improvement in robust clinical reasoning.
Motivation
Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often operating on idealized, clean EHRs via static SQL generation rather than interactive execution.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.