Benchmark Radar
AI BENCHMARK PROFILE

JudgmentBench

General AIKnowledge & Reasoning

A dataset of 30 legal tasks with rubric scores and pairwise preference judgments from practicing attorneys, used to compare rubric-based scoring and comparative judgment.

Released
2026-05-24
Readiness
Paper only
Primary field
General AI

Why it matters

Supports research on expert judgment elicitation and aggregation in domains without ground truth.

Motivation

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.