AI BENCHMARK PROFILE
DrugBench
DrugBench evaluates AI control protocols for mitigating medication-related harm using 3,671 medical conversations and FDA drug labels, covering drug interactions, contraindications, dosing constraints, and patient action restrictions. Introduces severity-based monitoring.
- Released
- 2026-06-10
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Addresses the safety-critical need to evaluate external safeguards for LLMs in medical QA, beyond simple accuracy. Useful for developers of safe medical AI systems and control protocols.
Motivation
Large Language Models have the potential to expand and improve the access to clinical information by enabling new ways of interacting with medical knowledge in natural language.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.