Benchmark Radar
AI BENCHMARK PROFILE

MedReaMM

Health & Life SciencesMultimodal Perception

Evaluates LLMs on multimodal clinical diagnostic synthesis using 625 expert-validated cases with medical images and ICD-11 diagnoses.

Released
2026-08-23
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Clinical diagnosis requires integrating heterogeneous evidence, and this benchmark highlights the gap between current models and expert-level synthesis.

Motivation

The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.