Benchmark Radar
AI BENCHMARK PROFILE

SPIKE-Bench

Health & Life SciencesSafety & Trustworthiness

SPIKE-Bench evaluates LLM biosecurity risks using 631 curated toxin-design prompts and a three-stage funnel (compliance, plausibility, predicted toxicity) producing the Functional Harmfulness Rate.

Released
2026-08-03
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Provides a function-aware metric for biosecurity evaluation beyond refusal rates, enabling comparison of models on predicted functional harm.

Motivation

Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.