SafePyramid
SafePyramid evaluates in-context policy guardrailing across 1,000 multi-turn conversations and 3,000 application-specific policies containing 61,699 natural-language rules, organized into three hierarchical capability levels (L0, L1, L2). Scoring is based on violated-rule set prediction with metrics RMR and RDR, and the evaluation harness supports any API or local model.
- Released
- 2026-06-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
SafePyramid addresses the gap in evaluating guardrails under application-specific policies rather than fixed taxonomies, providing a structured test for rule understanding, dependency resolution, and adaptation to novel frameworks. It enables direct comparison of frontier LLMs and configurable guardrails on a realistic safety task.
Motivation
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on predefined risk taxonomies.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.