Benchmark Radar
AI BENCHMARK PROFILE

ToolMenuBench

General AIKnowledge & Reasoning

ToolMenuBench is a benchmark for evaluating tool-menu filtering strategies in multi-step LLM agents. It varies tool-menu size, distractor type, state-dependent structure, and risk exposure, and reports filter-level and downstream metrics such as task success, tool calls, and token usage.

Released
2026-06-13
Readiness
Paper only
Primary field
General AI

Why it matters

ToolMenuBench addresses the gap in evaluating how tool-menu construction affects reliability, efficiency, and risk in tool-augmented agents. It provides a reusable framework for studying the agent-interface problem with controlled settings.

Motivation

Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether a model can call a tool correctly rather than how the visible tool menu shapes reliability, efficiency, and safety-relevant risk exposure.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.