AI BENCHMARK PROFILE
REDACT
REDACT is a multilingual benchmark for PII detection with 13,427 records, 51 entity types, and controlled generation axes. It allows stratified evaluation via metadata fields and includes an evaluation harness.
- Released
- 2026-06-18
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
PII detection lacks controlled benchmarks. REDACT provides systematic variation and layered evaluation to reveal failure conditions, aiding detector robustness.
Motivation
Benchmark infrastructure for personally identifiable information (PII) detection remains limited: existing corpora cover few entity types, use ad hoc generation conditions, and do not show which surface conditions cause detector failures.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.