OpAI-Bench
OpAI-Bench evaluates AI-text detection across document, sentence, token, and span granularities using operation-guided progressive human-to-AI revision trajectories with nine versions per sample.
- Released
- 2026-06-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks focus on static outputs, missing progressive co-editing. OpAI-Bench provides a controlled testbed with multi-granularity provenance to analyze how AI-authorship signals emerge and accumulate, revealing non-monotonic detection patterns.
Motivation
As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.