Benchmark Radar
AI BENCHMARK PROFILE

GuardianAgentBench

General AIKnowledge & ReasoningGuardianAgentBench Team

GuardianAgentBench (GABench) is a benchmark of 580 agent scenarios across six domains, evaluated on three production frameworks (LangChain, LlamaIndex, Vectara), with multi-stage validation and five adversarial attack modes to assess tool-use correctness and safety.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

LLM agents need structured evaluation of tool selection and failure modes; GABench provides a reusable framework with multiple attack modes and guardrail assessment, offering a fine-grained view of where agents fail.

Motivation

As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.