Benchmark Radar
AI BENCHMARK PROFILE

VARM-Bench

General AIMultimodal Perception

VARM-Bench evaluates verifiable structured reasoning in Chinese abusive-speech moderation. It uses field-anchored chain-of-thought rationales with six decision fields, and a deterministic protocol assessing field correctness, alignment, output validity, and record errors.

Released
2026-08-16
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks support classification but not verifiable reasoning. VARM-Bench provides an auditable protocol for evaluating moderation rationales, revealing that strong label performance can conceal errors in complete records.

Motivation

The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.