AIR-BENCH
AIR-BENCH Live is a self-evolving safety benchmark for foundation models, with an automated pipeline that updates risk taxonomy and prompts based on new regulations. It evaluates model safety across multilingual prompts.
- Released
- 2026-07-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
This benchmark aims to keep pace with evolving AI risks and regulations, providing a dynamic evaluation tool for model safety. Its automated updates could help maintain relevance, but the lack of a fixed protocol limits comparability.
Motivation
Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.