Benchmark Radar
AI BENCHMARK PROFILE

OpAI-Bench

General AIMultimodal PerceptionVILA-Lab

OpAI-Bench evaluates AI-text detection across document, sentence, token, and span granularities using operation-guided progressive human-to-AI revision trajectories with nine versions per sample.

Released
2026-06-04
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on static outputs, missing progressive co-editing. OpAI-Bench provides a controlled testbed with multi-granularity provenance to analyze how AI-authorship signals emerge and accumulate, revealing non-monotonic detection patterns.

Motivation

As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.