AI BENCHMARK PROFILE
MedPRESS
MedPRESS evaluates LLM sycophancy in multi-turn medical dialogues, containing 600 five-turn scenarios across three families with structured judging and safety metrics.
- Released
- 2026-08-03
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Measures robustness under conversational pressure, a gap in static medical safety evaluations.
Motivation
Large language models (LLMs) are increasingly used for health-related advice.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.