Benchmark Radar
AI BENCHMARK PROFILE

ABC-Bench

Health & Life SciencesRobotics & Autonomous SystemsKnowledge & Reasoning

ABC-Bench evaluates LLM agents on biosecurity-relevant tasks including liquid handling robot code generation, DNA fragment design, and DNA synthesis screening evasion, with wet-lab validation.

Released
2026-06-09
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Measures agentic AI capabilities relevant to biosecurity, offering a standardized protocol to assess dual-use risks and inform safeguards in biological research contexts.

Motivation

Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.