AI BENCHMARK PROFILE
CLIR-Bench
A benchmark for question answering over irregular clinical time series from ICU records, containing 6,600 QA instances across 11 clinical variables and 11 tasks. It evaluates answer accuracy and evidence use through explicit temporal evidence and answer derivation rules.
- Released
- 2026-07-10
- Readiness
- Inspectable
- Primary field
- Health & Life Sciences
Why it matters
Existing benchmarks focus on regular time-series or static medical QA, while real ICU data is sparse and asynchronous. This benchmark provides a way to assess whether models can reason over irregular temporal evidence, addressing a gap in clinical NLP evaluation.
Motivation
Clinical time series are central to patient monitoring, risk assessment, and clinical decision support.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.