GuardianAgentBench
GuardianAgentBench (GABench) is a benchmark of 580 agent scenarios across six domains, evaluated on three production frameworks (LangChain, LlamaIndex, Vectara), with multi-stage validation and five adversarial attack modes to assess tool-use correctness and safety.
- Released
- 2026-07-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
LLM agents need structured evaluation of tool selection and failure modes; GABench provides a reusable framework with multiple attack modes and guardrail assessment, offering a fine-grained view of where agents fail.
Motivation
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.