Benchmark Radar
AI BENCHMARK PROFILE

AgentJudgeBench

General AIKnowledge & Reasoning

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined.

Released
2026-08-27
Readiness
Paper only
Primary field
General AI

Motivation

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.