Benchmark Radar
AI BENCHMARK PROFILE

REDAgentBench

CybersecurityAgents

Evaluates LLM agent safety through executable red-teaming, adversarial case generation, and verification of harmful effects across 1,661 cases and five service surfaces.

Released
2026-08-11
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Provides an executable and measurable approach for agent safety evaluation beyond aggregate attack success rates.

Motivation

Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.