Benchmark Radar
AI BENCHMARK PROFILE

PathoArgus-Bench

Health & Life SciencesMultimodal PerceptionLong Context & MemoryPathoArgus-Bench Team

PathoArgus-Bench evaluates evidence-grounded visual reasoning in whole-slide pathology. It comprises 22,078 multiple-choice questions from 4,913 patients across 15 TCGA projects, testing availability, accessibility, use, and responsiveness of evidence under a fixed reader budget.

Released
2026-08-18
Readiness
Inspectable
Primary field
Health & Life Sciences

Why it matters

Addresses the gap where final answer accuracy is insufficient to establish evidence grounding in pathology AI. Provides a protocol to assess whether models truly use supplied tissue evidence, offering a more rigorous evaluation for clinical deployment.

Motivation

Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric vulnerable to linguistic priors and benchmark regularities, and insufficient to establish that predictions are grounded in the supplied tissue.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.