AutoMedBench
AutoMedBench evaluates autonomous AI agents on end-to-end medical-AI research tasks spanning segmentation, image enhancement, VQA, report generation, and lesion detection. Tasks follow a five-stage workflow (Plan, Setup, Validate, Inference, Submit), with scoring based on both final task performance and stage-level rubric scores.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Existing medical agent benchmarks focus on final outputs, obscuring failure points. AutoMedBench provides granular stage-level scoring and error analysis, enabling targeted assessment of agent capabilities and highlighting bottlenecks like verification and submission, which can guide improvement priorities in automated medical research systems.
Motivation
Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form clinical question answering.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.