Benchmark Radar
AI BENCHMARK PROFILE

ToolRobustBench

General AICoding & Software Engineering

Evaluates tool-calling agents under perturbations across the tool-use pipeline, attributing failures to selection, grounding, argument binding, and feedback handling.

Released
2026-08-23
Readiness
Paper only
Primary field
General AI

Why it matters

Provides deterministic, cascade-aware diagnosis of robustness beyond clean tool-calling accuracy.

Motivation

Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to invoke external systems and complete tasks beyond text generation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.