Benchmark Radar
AI BENCHMARK PROFILE

MedPIC-Bench

CybersecurityHealth & Life SciencesSafety & Trustworthiness

MedPIC-Bench evaluates patient-specific medication-safety reasoning in LLMs using guideline-following and paired counterfactual questions, with 467 questions annotated across six dimensions.

Released
2026-08-04
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Static medication-safety accuracy may not reflect whether models use patient information to determine rule applicability. This benchmark aims to make conditional rule application measurable.

Motivation

Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.