Benchmark Radar
AI BENCHMARK PROFILE

EvalAwareBench

General AISafety & Trustworthiness

EvalAwareBench is a factor-controlled benchmark of 100 paired safety-capability tasks, each toggling one of eight evaluation-awareness triggers while holding the underlying request fixed. It measures recognition and behavioral change in language models under evaluation.

Released
2026-05-21
Readiness
Paper only
Primary field
General AI

Why it matters

Evaluation awareness can distort benchmark results, especially for safety evaluations. EvalAwareBench enables controlled measurement of model sensitivity to evaluative signals, helping researchers identify and mitigate threats to benchmark validity.

Motivation

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.