SafeClawBench
SafeClawBench is a staged benchmark for tool-using LLM agent security with 600 adversarial tasks across six attack families, reporting three endpoints: semantic attack acceptance, audit-visible harm evidence, and sandbox-observed tool/state harm.
- Released
- 2026-06-16
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
SafeClawBench addresses the evaluation gap where existing benchmarks collapse distinct security failure stages into a single metric, making it hard to distinguish semantic compliance from actual harm. It provides separate, comparable scores across models and prompt policies, aiding in selecting agents and defenses based on the specific type of security risk.
Motivation
Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.