AI BENCHMARK PROFILE
MedPIC-Bench
MedPIC-Bench evaluates patient-specific medication-safety reasoning in LLMs using guideline-following and paired counterfactual questions, with 467 questions annotated across six dimensions.
- Released
- 2026-08-04
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Static medication-safety accuracy may not reflect whether models use patient information to determine rule applicability. This benchmark aims to make conditional rule application measurable.
Motivation
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.