Benchmark Radar
AI BENCHMARK PROFILE

XIH-Bench

General AIKnowledge & Reasoningg1moon

Evaluates instruction hierarchy compliance in multilingual LLMs using same-language and cross-language conflicts across six languages, four domains (rule-following, safety, task-execution, persona), and three hierarchy types (system-user, system-tool, user-tool).

Released
2026-07-26
Readiness
Runnable
Primary field
General AI

Why it matters

Existing instruction hierarchy benchmarks are largely English-centric, leaving a gap in assessing multilingual safety and reliability. This benchmark provides a reusable protocol to measure how language choice affects model compliance, supporting safer deployment in multilingual contexts.

Motivation

Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.