ClinOCR-Bench
ClinOCR-Bench is a public dataset of 384 scanned clinical documents across six artifact subsets (normal, handwriting, poor quality, rotation, tables, mixed). It evaluates OCR systems on clinical text extraction, with ground truth transcriptions and template-aware train/test splits supporting 0-shot and 1-shot evaluation.
- Released
- 2026-07-04
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Existing clinical OCR evaluations rely on private data lacking common scan artifacts. ClinOCR-Bench provides a standardized, realistic set of clinical documents with controlled artifacts, enabling reproducible comparison of OCR and vision-language models on real-world clinical scanning challenges.
Motivation
Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.