Benchmark Radar
AI BENCHMARK PROFILE

PathAgentBench

Health & Life SciencesMultimodal PerceptionSearch & Retrieval

PathAgentBench evaluates vision-language models on whole-slide pathology images across four capabilities: image-to-text matching, text-to-image retrieval, diagnostic-region localization, and multi-scale reasoning. It includes 1,822 TCGA WSIs and 17,135 diagnostic paths.

Released
2026-07-21
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Most pathology benchmarks use pre-cropped patches, not whole-slide exploration. PathAgentBench provides a unified framework with annotated paths, revealing a significant gap in evidence acquisition and supporting progress in autonomous WSI diagnosis.

Motivation

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.