Benchmark Radar
AI BENCHMARK PROFILE

BabelJudge

General AISafety & Trustworthiness

BabelJudge audits LLM-as-a-judge reliability across languages and agent trajectories, measuring position bias, verbosity bias, order inconsistency, and cross-lingual degradation without human labels, providing a composite bias-penalised reliability score.

Released
2026-06-21
Readiness
Runnable
Primary field
General AI

Why it matters

It addresses the gap that raw accuracy hides systematic judge biases, offering a standardized method to quantify reliability for automated evaluation, crucial for trustworthy model comparisons and training data quality in multilingual and agentic contexts.

Motivation

LLM-as-a-judge has become the dominant approach to scalable evaluation in NLP pipelines, yet judges themselves carry systematic biases that raw accuracy hides: they favor responses placed in slot A (position bias), they prefer longer responses regardless of quality (verbosity bias), and their reliability degrades sharply in lower-resource languages.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.