SCR-Bench
SCR-Bench is a benchmark for evaluating security risks in composed LLM agent skill workflows. It includes three sub-benchmarks (SCR-CapFlow, SCR-TrustLift, SCR-AuthBlur) that measure attack success rates, trust transfer, and authorization confusion in sandboxed environments.
- Released
- 2026-06-13
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
SCR-Bench addresses the gap in evaluating agent skill security at the path level, where skills benign in isolation become harmful in composition. It provides a controlled, sandboxed environment for assessing composition-induced risks.
Motivation
Skills are becoming the capability layer through which LLM agents turn plans into actions, but their use introduces security risks such as data leakage, unauthorized operations, and tool misuse.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.