AI BENCHMARK PROFILE
ToolFailBench
ToolFailBench evaluates tool-use failures in LLM agents across 1,000 tasks in finance, medicine, law, cybersecurity, and real estate. It labels traces with Tool-Skip, Result-Ignore, Output-Fabrication, and Unnecessary-Tool-Use, using a rule classifier and LLM judges.
- Released
- 2026-07-06
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Aggregate accuracy hides distinct failure modes in tool use. This benchmark separates models that fail to call tools from those that call but ignore results, enabling targeted diagnosis of agent reliability.
Motivation
Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.