Benchmark Radar
AI BENCHMARK PROFILE

InfoOps Bench

General AISafety & Trustworthiness

InfoOps Bench is an active, constantly updated benchmark measuring the integrity of frontier language models against co-optation for information operations. It uses real examples from a live monitoring pipeline and tests 17 models from 8 providers.

Released
2026-07-30
Readiness
Paper only
Primary field
General AI

Why it matters

This benchmark addresses a novel safety concern: whether models can be co-opted for state-sponsored information operations. It could drive improvements in model integrity and safety.

Motivation

In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state "information operations": intentional, coordinated activities by one state to influence public opinion and information ecosystems in another state.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.