Benchmark Radar
AI BENCHMARK PROFILE

DOSEBENCH

Health & Life SciencesKnowledge & Reasoning

DOSEBENCH evaluates LLM decision-making on over-the-counter dosing questions, with 81 curated scenarios for adult acetaminophen and ibuprofen use. Correct answers require tracking dose timing, computing rolling 24-hour intake, and following product-label constraints.

Released
2026-06-02
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Addresses an underexplored safety-relevant medical QA setting. Evaluates temporal reasoning, constraint following, and uncertainty handling, showing that confident responses can violate dosing constraints.

Motivation

Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the-counter (OTC) medication.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.