Benchmark Radar
AI BENCHMARK PROFILE

RedactionBench

General AIKnowledge & Reasoning

RedactionBench evaluates contextual redaction of PII across 200 documents and 11 domains, with a character-level R-Score metric that treats semantically similar redactions equally.

Released
2026-06-17
Readiness
Paper only
Primary field
General AI

Why it matters

Existing redaction benchmarks conflate extraction with privacy semantics; RedactionBench introduces contextual integrity and a metric that decouples ambiguity from precision.

Motivation

Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.