ToxScreen
A benchmark of roughly 800 backdoored LLMs across attack objectives, trigger mechanisms, poisoning rates, model scales, and training mechanisms. Evaluates whether a defender can recover a planted trigger given white-box access and behavior of concern.
- Released
- 2026-07-29
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Backdoor recovery in LLMs is a critical security challenge; this benchmark provides a standardized testbed for comparing trigger-recovery methods under realistic constraints, informing practical defense strategies.
Motivation
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.