Benchmark Radar
AI BENCHMARK PROFILE

ToolFailBench

CybersecurityFinance & EconomicsKnowledge & Reasoning

ToolFailBench evaluates tool-use failures in LLM agents across 1,000 tasks in finance, medicine, law, cybersecurity, and real estate. It labels traces with Tool-Skip, Result-Ignore, Output-Fabrication, and Unnecessary-Tool-Use, using a rule classifier and LLM judges.

Released
2026-07-06
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Aggregate accuracy hides distinct failure modes in tool use. This benchmark separates models that fail to call tools from those that call but ignore results, enabling targeted diagnosis of agent reliability.

Motivation

Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.