MedCTA
MedCTA evaluates clinical tool agents on 107 clinician-verified, step-implicit tasks with multimodal inputs (radiology images, pathology slides, reports) and 5 deployed tools. Metrics cover tool selection, argument validity, execution stability, trajectory fidelity, and outcome quality.
- Released
- 2026-06-10
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Fills a gap in medical AI evaluation by going beyond single-turn QA to test agentic tool use, planning, and reliability in real-world clinical workflows. Useful for auditing and advancing trustworthy medical AI agents.
Motivation
To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.