Benchmark Radar
AI BENCHMARK PROFILE

BenchBench-Protocol

General AIKnowledge & Reasoning

Evaluates LLMs on 149 protocol-modification tasks reconstructed from real changes scientists made to published wet-lab protocols, using weighted rubric scoring.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

Assesses routine wet-lab adaptation reasoning from real experimental modifications, filling a gap in grounded life-science evaluation beyond expert-elicited tasks.

Motivation

We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.