Benchmark Radar
AI BENCHMARK PROFILE

RuVerBench

General AIKnowledge & ReasoningTsinghua University

RuVerBench evaluates LLM-as-a-judge reliability for rubric verification in agentic scenarios. It includes 2,458 instances across deep research and agentic coding, each with a model-generated output, a rubric, and a human-annotated label indicating rubric satisfaction.

Released
2026-06-29
Readiness
Runnable
Primary field
General AI

Why it matters

Rubric-based scoring with LLM judges is common but under-validated, especially for agentic outputs. RuVerBench provides a reusable benchmark to compare judge models and strategies, enabling decisions on model selection and scoring protocol.

Motivation

Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.