VanillaBench
VanillaBench evaluates the clean accuracy gap between adversarially trained models and vanilla (non-robust) reference models across four threat models. It defines a protocol for comparing robustness-accuracy trade-offs using model accuracy on standard benchmarks.
- Released
- 2026-07-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses the evaluation gap in adversarial robustness research by quantifying the cost of robustness, providing practitioners with information needed to make informed deployment decisions regarding accuracy versus robustness.
Motivation
Adversarial robustness research has produced hundreds of defended models over the past decade, yet the literature almost universally reports robustness results in isolation: standard (clean) accuracy and adversarial accuracy of the robust model are shown, but the gap to the corresponding vanilla model is rarely quantified.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.