Benchmark Radar
AI BENCHMARK PROFILE

LongJudgeBench

General AIKnowledge & ReasoningLongJudgeBench Team

LongJudgeBench evaluates LLM-as-a-judge performance on long-form outputs across six datasets covering pointwise, pairwise, and listwise protocols, with bilingual tasks and multiple prompt variants. It measures agreement with human judgments using accuracy, Spearman, and Kendall's tau.

Released
2026-06-01
Readiness
Runnable
Primary field
General AI

Why it matters

Existing meta-evaluation benchmarks focus on short-form outputs, leaving a gap for long-form evaluation. This benchmark provides a standardized way to assess judge reliability across diverse scenarios, helping practitioners select or develop judges for long-form tasks where current models show instability.

Motivation

As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.