Benchmark Radar
AI BENCHMARK PROFILE

PluginEval

General AIKnowledge & Reasoning

PluginEval evaluates tool routing in LLMs via three decision types (missed, spurious, parameter errors) across difficulty levels, using deterministic validation and real API execution for reliable signals.

Released
2026-08-09
Readiness
Paper only
Primary field
General AI

Why it matters

It overcomes limitations of power-law data distributions and unvalidated LLM judgments, offering a diagnostic error profile for function calling and enabling more reliable agent evaluation.

Motivation

Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.