Benchmark Radar
AI BENCHMARK PROFILE

RoleCDE

General AIKnowledge & ReasoningRoleCDE Team

RoleCDE evaluates role-playing agents under structured conflicts between role-specific values and alignment-oriented constraints. It comprises approximately 8,000 role profiles and 240,000 dilemma instances across three difficulty levels and eight role categories, with scoring via LLM-as-a-judge.

Released
2026-06-01
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on surface fidelity and lack coverage of decision-making under role-alignment value conflicts. RoleCDE provides a systematic evaluation of how agents resolve such conflicts, revealing systematic biases and offering a tool for improving alignment and role consistency.

Motivation

Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision making under role-alignment value conflicts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.