LongJudgeBench
LongJudgeBench evaluates LLM-as-a-judge performance on long-form outputs across six datasets covering pointwise, pairwise, and listwise protocols, with bilingual tasks and multiple prompt variants. It measures agreement with human judgments using accuracy, Spearman, and Kendall's tau.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing meta-evaluation benchmarks focus on short-form outputs, leaving a gap for long-form evaluation. This benchmark provides a standardized way to assess judge reliability across diverse scenarios, helping practitioners select or develop judges for long-form tasks where current models show instability.
Motivation
As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.