AI BENCHMARK PROFILE
EvoHarmBench
Dynamic adversarial evaluation framework that evolves evasion strategies across 229 semantic sub-clusters from 5,002 real-world samples to assess content moderation systems.
- Released
- 2026-08-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the gap between static benchmarks and online moderation by providing an iterative evaluation protocol with public code promised, supporting more robust safety testing.
Motivation
Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content platforms where users continuously revise their expressions in response to moderation feedback.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.