Benchmark Radar
AI BENCHMARK PROFILE

MedPRESS

Health & Life SciencesKnowledge & Reasoning

MedPRESS evaluates LLM sycophancy in multi-turn medical dialogues, containing 600 five-turn scenarios across three families with structured judging and safety metrics.

Released
2026-08-03
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Measures robustness under conversational pressure, a gap in static medical safety evaluations.

Motivation

Large language models (LLMs) are increasingly used for health-related advice.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.