Benchmark Radar
AI BENCHMARK PROFILE

MedCTA

Health & Life SciencesMultimodal PerceptionIVUL-KAUST

MedCTA evaluates clinical tool agents on 107 clinician-verified, step-implicit tasks with multimodal inputs (radiology images, pathology slides, reports) and 5 deployed tools. Metrics cover tool selection, argument validity, execution stability, trajectory fidelity, and outcome quality.

Released
2026-06-10
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Fills a gap in medical AI evaluation by going beyond single-turn QA to test agentic tool use, planning, and reliability in real-world clinical workflows. Useful for auditing and advancing trustworthy medical AI agents.

Motivation

To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.