BabelJudge
BabelJudge audits LLM-as-a-judge reliability across languages and agent trajectories, measuring position bias, verbosity bias, order inconsistency, and cross-lingual degradation without human labels, providing a composite bias-penalised reliability score.
- Released
- 2026-06-21
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It addresses the gap that raw accuracy hides systematic judge biases, offering a standardized method to quantify reliability for automated evaluation, crucial for trustworthy model comparisons and training data quality in multilingual and agentic contexts.
Motivation
LLM-as-a-judge has become the dominant approach to scalable evaluation in NLP pipelines, yet judges themselves carry systematic biases that raw accuracy hides: they favor responses placed in slot A (position bias), they prefer longer responses regardless of quality (verbosity bias), and their reliability degrades sharply in lower-resource languages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.