AI BENCHMARK PROFILE
SPIKE-Bench
SPIKE-Bench evaluates LLM biosecurity risks using 631 curated toxin-design prompts and a three-stage funnel (compliance, plausibility, predicted toxicity) producing the Functional Harmfulness Rate.
- Released
- 2026-08-03
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Provides a function-aware metric for biosecurity evaluation beyond refusal rates, enabling comparison of models on predicted functional harm.
Motivation
Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.