Benchmark Radar
AI BENCHMARK PROFILE

ClinLens

Health & Life SciencesMultimodal Perception

ClinLens evaluates clinical data-science agents on 200 executable tasks over five linked MIMIC resources (EHR, notes, ECG, chest X-rays, echocardiograms), organized by a 4x5 taxonomy of patient-time scopes and analysis capabilities. Scoring uses a STRICTPASS metric requiring correct artifacts, semantics, and final answers.

Released
2026-07-28
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Existing benchmarks isolate medical QA or table reasoning, lacking integrated longitudinal clinical data science. ClinLens fills this gap with a program-first reverse synthesis approach, exposing a gap between runnable code and correct clinical analysis, guiding improvements in clinical agent reliability.

Motivation

Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.