ToolMenuBench
ToolMenuBench is a benchmark for evaluating tool-menu filtering strategies in multi-step LLM agents. It varies tool-menu size, distractor type, state-dependent structure, and risk exposure, and reports filter-level and downstream metrics such as task success, tool calls, and token usage.
- Released
- 2026-06-13
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
ToolMenuBench addresses the gap in evaluating how tool-menu construction affects reliability, efficiency, and risk in tool-augmented agents. It provides a reusable framework for studying the agent-interface problem with controlled settings.
Motivation
Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether a model can call a tool correctly rather than how the visible tool menu shapes reliability, efficiency, and safety-relevant risk exposure.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.