Benchmark Radar
AI BENCHMARK PROFILE

IH-Benchmark

General AISafety & Trustworthiness

IH-Benchmark evaluates instruction-hierarchy robustness in LLMs via conflicting instructions from system, user, and tool outputs, covering 44 constraint families across five domains with a binary pass/fail protocol.

Released
2026-07-28
Readiness
Paper only
Primary field
General AI

Why it matters

Existing instruction-hierarchy benchmarks cover limited conflict types and tool interactions. IH-Benchmark provides a systematic evaluation across conflict surfaces, constraint types, and attack presentations, revealing that robustness is not a single capability but a set of behaviors with distinct failure modes.

Motivation

When a language model receives conflicting instructions from different priority levels, which one does it actually follow?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.