Benchmark Radar
AI BENCHMARK PROFILE

SafePyramid

General AISafety & TrustworthinessByteDance

SafePyramid evaluates in-context policy guardrailing across 1,000 multi-turn conversations and 3,000 application-specific policies containing 61,699 natural-language rules, organized into three hierarchical capability levels (L0, L1, L2). Scoring is based on violated-rule set prediction with metrics RMR and RDR, and the evaluation harness supports any API or local model.

Released
2026-06-29
Readiness
Runnable
Primary field
General AI

Why it matters

SafePyramid addresses the gap in evaluating guardrails under application-specific policies rather than fixed taxonomies, providing a structured test for rule understanding, dependency resolution, and adaptation to novel frameworks. It enables direct comparison of frontier LLMs and configurable guardrails on a realistic safety task.

Motivation

In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on predefined risk taxonomies.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.