OpenSafeIntent
OpenSafeIntent is a benchmark of controlled prompt-sets that vary user intent while holding the underlying task fixed. Each datapoint contains benign, dual-use, and malicious variants of the same task, and models are evaluated on whether they calibrate assistance across intent shifts rather than only appearing safe on average.
- Released
- 2026-07-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It evaluates safety as intent-calibrated behavior over controlled task variants, addressing limitations of evaluating safety on isolated prompts and providing a way to assess safe completion across subtle intent changes.
Motivation
Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.