Benchmark Radar
AI BENCHMARK PROFILE

PolicyShiftBench

General AISafety & Trustworthiness

PolicyShiftBench evaluates policy-adaptive image guardrailing: given an image and a current policy, a model must output a pass/block decision plus optional violated category IDs. The benchmark comprises 2,000 policy-discriminative instances over 265 images, each paired with multiple policy-conditioned prompts. Scoring uses binary pass/block accuracy and category attribution metrics.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Existing image safety benchmarks assume safety is a fixed property of an image, whereas real deployments vary policies across products and regions. PolicyShiftBench measures whether models can bind image evidence to the active policy rather than relying on image-level priors, providing a practical evaluation for guardrails in dynamic policy environments.

Motivation

Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.