LongMedBench
A benchmark for long-horizon clinical decision-making using EHR data from MIMIC-IV, comprising 335 patients with multi-session interactions and three evaluation suites: fact-based QA, temporal reasoning, and long-horizon decision-making.
- Released
- 2026-07-10
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Current medical agent evaluations emphasize short-context tasks, while real clinical care requires aggregating evidence over extended periods. This benchmark addresses the need for realistic long-horizon assessment, but the paper does not specify if the benchmark is publicly available for reuse or ongoing submission.
Motivation
In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.