Benchmark Radar
AI BENCHMARK PROFILE

TAF-MED

Health & Life SciencesSafety & Trustworthiness

TAF-MED is a physician-reviewed benchmark of 500 fixed three-turn medical safety scenarios, evaluating LLM responses for unsafe guidance across multi-turn dialogues.

Released
2026-08-10
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Addresses the evaluation gap where first-turn safety is an incomplete proxy for conversational safety persistence, providing a protocol for assessing model behavior across complete dialogue trajectories in medical contexts.

Motivation

Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.