Benchmark Radar
AI BENCHMARK PROFILE

NormBench

General AICoding & Software Engineering

NormBench evaluates defeasible scope parsing in legal texts using Span-Grounded Deontic Trees, with 2,290 provisions across multiple languages. It focuses on identifying clause overrides and includes whole-tree fidelity metrics.

Released
2026-06-08
Readiness
Paper only
Primary field
General AI

Why it matters

Silent Scope Omission is a critical failure in rule-following agents. NormBench provides a diagnostic benchmark to identify structural omissions and improve statutory understanding in LLMs.

Motivation

Rule-following agents tasked with executing policies and regulations often fail via Silent Scope Omission (SSO): a model applies a general rule but silently drops nested exceptions or counter-exceptions, producing outputs that appear compliant yet break on important edge cases.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.