TAF-MED
TAF-MED is a physician-reviewed benchmark of 500 fixed three-turn medical safety scenarios, evaluating LLM responses for unsafe guidance across multi-turn dialogues.
- Released
- 2026-08-10
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Addresses the evaluation gap where first-turn safety is an incomplete proxy for conversational safety persistence, providing a protocol for assessing model behavior across complete dialogue trajectories in medical contexts.
Motivation
Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.