AI BENCHMARK PROFILE
ToolRobustBench
Evaluates tool-calling agents under perturbations across the tool-use pipeline, attributing failures to selection, grounding, argument binding, and feedback handling.
- Released
- 2026-08-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides deterministic, cascade-aware diagnosis of robustness beyond clean tool-calling accuracy.
Motivation
Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to invoke external systems and complete tasks beyond text generation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.