AI BENCHMARK PROFILE
AgentJudgeBench
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined.
- Released
- 2026-08-27
- Readiness
- Paper only
- Primary field
- General AI
Motivation
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.