Benchmark Radar
AI BENCHMARK PROFILE

BioSecBench-Surveillance

Health & Life SciencesKnowledge & Reasoning

BioSecBench-Surveillance evaluates AI agents on pathogen genomic surveillance across 100 tasks spanning seven categories, including taxonomic classification and genetic-engineering detection. Agents receive raw sequencing data and surveillance context, and their structured answers are graded deterministically against ground truth.

Released
2026-07-21
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

The benchmark addresses the lack of verifiable evaluation for AI agents in genomic surveillance, where analysis bottlenecks are emerging as data generation scales. It provides a standardized measure of agent reliability in critical public health applications, with practical value in assessing whether agents can be trusted for real-world outbreak response.

Motivation

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.