AI BENCHMARK PROFILE
IMCBench
IMCBench evaluates multimodal LLMs in image-grounded, multi-turn medical conversations, scoring safety, accuracy, and uncertainty use on a 1-5 scale.
- Released
- 2026-06-26
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Addresses the gap in medical AI benchmarks by combining clinical images with multi-turn dialogue, enabling assessment of diagnostic accuracy alongside patient safety and uncertainty management.
Motivation
Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.