IH-Benchmark
IH-Benchmark evaluates instruction-hierarchy robustness in LLMs via conflicting instructions from system, user, and tool outputs, covering 44 constraint families across five domains with a binary pass/fail protocol.
- Released
- 2026-07-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing instruction-hierarchy benchmarks cover limited conflict types and tool interactions. IH-Benchmark provides a systematic evaluation across conflict surfaces, constraint types, and attack presentations, revealing that robustness is not a single capability but a set of behaviors with distinct failure modes.
Motivation
When a language model receives conflicting instructions from different priority levels, which one does it actually follow?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.