ClinLens
ClinLens evaluates clinical data-science agents on 200 executable tasks over five linked MIMIC resources (EHR, notes, ECG, chest X-rays, echocardiograms), organized by a 4x5 taxonomy of patient-time scopes and analysis capabilities. Scoring uses a STRICTPASS metric requiring correct artifacts, semantics, and final answers.
- Released
- 2026-07-28
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Existing benchmarks isolate medical QA or table reasoning, lacking integrated longitudinal clinical data science. ClinLens fills this gap with a program-first reverse synthesis approach, exposing a gap between runnable code and correct clinical analysis, guiding improvements in clinical agent reliability.
Motivation
Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.