BioSecBench-Surveillance
BioSecBench-Surveillance evaluates AI agents on pathogen genomic surveillance across 100 tasks spanning seven categories, including taxonomic classification and genetic-engineering detection. Agents receive raw sequencing data and surveillance context, and their structured answers are graded deterministically against ground truth.
- Released
- 2026-07-21
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
The benchmark addresses the lack of verifiable evaluation for AI agents in genomic surveillance, where analysis bottlenecks are emerging as data generation scales. It provides a standardized measure of agent reliability in critical public health applications, with practical value in assessing whether agents can be trusted for real-world outbreak response.
Motivation
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.