AI BENCHMARK PROFILE
BenchBench-Protocol
Evaluates LLMs on 149 protocol-modification tasks reconstructed from real changes scientists made to published wet-lab protocols, using weighted rubric scoring.
- Released
- 2026-08-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Assesses routine wet-lab adaptation reasoning from real experimental modifications, filling a gap in grounded life-science evaluation beyond expert-elicited tasks.
Motivation
We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.