Benchmark Radar
AI BENCHMARK PROFILE

LongMedBench

Health & Life SciencesKnowledge & Reasoning

A benchmark for long-horizon clinical decision-making using EHR data from MIMIC-IV, comprising 335 patients with multi-session interactions and three evaluation suites: fact-based QA, temporal reasoning, and long-horizon decision-making.

Released
2026-07-10
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Current medical agent evaluations emphasize short-context tasks, while real clinical care requires aggregating evidence over extended periods. This benchmark addresses the need for realistic long-horizon assessment, but the paper does not specify if the benchmark is publicly available for reuse or ongoing submission.

Motivation

In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.