Benchmark Radar
AI BENCHMARK PROFILE

PathMCQA

Health & Life SciencesMultimodal Perception

PathMMU is a massive multimodal expert-level benchmark for understanding and reasoning in pathology, containing 33,428 multimodal multi-choice questions and 24,067 images validated by seven pathologists. It evaluates Large Multimodal Models (LMMs) performance on pathology tasks, with the top-performing model GPT-4V achieving only 49.8% zero-shot performance compared to 71.8% for human pathologists.

Released
Unknown
Readiness
Paper only
Primary field
Health & Life Sciences

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.