Benchmark Radar
AI BENCHMARK PROFILE

HealthAgentBench

Health & Life SciencesAgentsMicrosoft

Evaluates AI agents on 54 realistic healthcare tasks across 7 categories, including medical imaging, EHR analysis, and clinical trial matching. Agents operate in terminal environments with task-specific verifiers and a final task success rate.

Released
2026-06-30
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Provides a standardized, realistic evaluation for agentic healthcare AI, revealing performance gaps across task types and informing model selection for clinical applications.

Motivation

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.