AI BENCHMARK PROFILE
PluginEval
PluginEval evaluates tool routing in LLMs via three decision types (missed, spurious, parameter errors) across difficulty levels, using deterministic validation and real API execution for reliable signals.
- Released
- 2026-08-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It overcomes limitations of power-law data distributions and unvalidated LLM judgments, offering a diagnostic error profile for function calling and enabling more reliable agent evaluation.
Motivation
Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.