CIAware-Bench
CIAware-Bench measures control intervention awareness in language models through four task domains: essay writing, BigCodeBench, Bash Arena, and SHADE-Arena. Models are tested on distinguishing their own trajectories from those modified by a control protocol, with variations in watermarking, side-task presence, and protocol.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
This benchmark addresses the evaluation gap of determining whether models can detect modifications to their trajectories, which is critical for AI control protocols. It provides practical value by informing the design of interventions that are harder for models to detect, thereby improving the robustness of AI oversight systems.
Motivation
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.