AI BENCHMARK PROFILE
HealthAgentBench
Evaluates AI agents on 54 realistic healthcare tasks across 7 categories, including medical imaging, EHR analysis, and clinical trial matching. Agents operate in terminal environments with task-specific verifiers and a final task success rate.
- Released
- 2026-06-30
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Provides a standardized, realistic evaluation for agentic healthcare AI, revealing performance gaps across task types and informing model selection for clinical applications.
Motivation
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.