Benchmark Radar
AI BENCHMARK PROFILE

REDACT

General AIMultimodal PerceptionREDACT Team

REDACT is a multilingual benchmark for PII detection with 13,427 records, 51 entity types, and controlled generation axes. It allows stratified evaluation via metadata fields and includes an evaluation harness.

Released
2026-06-18
Readiness
Paper only
Primary field
General AI

Why it matters

PII detection lacks controlled benchmarks. REDACT provides systematic variation and layered evaluation to reveal failure conditions, aiding detector robustness.

Motivation

Benchmark infrastructure for personally identifiable information (PII) detection remains limited: existing corpora cover few entity types, use ad hoc generation conditions, and do not show which surface conditions cause detector failures.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.