AI BENCHMARK PROFILE
MedReaMM
Evaluates LLMs on multimodal clinical diagnostic synthesis using 625 expert-validated cases with medical images and ICD-11 diagnoses.
- Released
- 2026-08-23
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Clinical diagnosis requires integrating heterogeneous evidence, and this benchmark highlights the gap between current models and expert-level synthesis.
Motivation
The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.