Benchmark Radar
AI BENCHMARK PROFILE

ContractScrub

General AIKnowledge & Reasoning

Evaluates LLMs on legal contract scrubbing tasks using hand-crafted contracts covering error categories like defined term misuse and inconsistent language.

Released
2026-08-20
Readiness
Paper only
Primary field
General AI

Why it matters

Provides the first formal evaluation of contract scrubbing, revealing that frontier models perform surprisingly poorly on this specific legal review task despite strong general benchmarks.

Motivation

Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.