Benchmark Radar
AI BENCHMARK PROFILE

IMCBench

Health & Life SciencesMultimodal Perception

IMCBench evaluates multimodal LLMs in image-grounded, multi-turn medical conversations, scoring safety, accuracy, and uncertainty use on a 1-5 scale.

Released
2026-06-26
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Addresses the gap in medical AI benchmarks by combining clinical images with multi-turn dialogue, enabling assessment of diagnostic accuracy alongside patient safety and uncertainty management.

Motivation

Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.