Benchmark Radar
AI BENCHMARK PROFILE

AIR-BENCH

General AISafety & Trustworthiness

AIR-BENCH Live is a self-evolving safety benchmark for foundation models, with an automated pipeline that updates risk taxonomy and prompts based on new regulations. It evaluates model safety across multilingual prompts.

Released
2026-07-06
Readiness
Paper only
Primary field
General AI

Why it matters

This benchmark aims to keep pace with evolving AI risks and regulations, providing a dynamic evaluation tool for model safety. Its automated updates could help maintain relevance, but the lack of a fixed protocol limits comparability.

Motivation

Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.