Benchmark Radar
AI BENCHMARK PROFILE

BioSecBench-Refusal

General AIKnowledge & Reasoning

BioSecBench-Refusal evaluates AI agents on 61 Routine and 46 Red-Team biosecurity-related tasks, measuring refusal rates and risk identification across multiple model configurations.

Released
2026-07-06
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the need for benchmarks that quantify both capability and safety in agentic biosecurity, helping developers calibrate models to avoid over-refusal while still detecting threats.

Motivation

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.