PathoArgus-Bench
PathoArgus-Bench evaluates evidence-grounded visual reasoning in whole-slide pathology. It comprises 22,078 multiple-choice questions from 4,913 patients across 15 TCGA projects, testing availability, accessibility, use, and responsiveness of evidence under a fixed reader budget.
- Released
- 2026-08-18
- Readiness
- Inspectable
- Primary field
- Health & Life Sciences
Why it matters
Addresses the gap where final answer accuracy is insufficient to establish evidence grounding in pathology AI. Provides a protocol to assess whether models truly use supplied tissue evidence, offering a more rigorous evaluation for clinical deployment.
Motivation
Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric vulnerable to linguistic priors and benchmark regularities, and insufficient to establish that predictions are grounded in the supplied tissue.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.