Benchmark Radar
AI BENCHMARK PROFILE

ParamBench

General AIKnowledge & Reasoning

ParamBench is a benchmark for evaluating LLM tool call parameter generation, built from real cloud-network APIs with difficulty tiers and exact match metrics. It is used to evaluate the proposed probe-guided training framework.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

Tool call parameter correctness is critical for execution yet understudied. ParamBench provides a systematic evaluation of parameter generation across difficulty levels, showing large improvements from probe-guided methods.

Motivation

Large language model agents derive much of their capability from tool use.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.