SkillVetBench
SkillVetBench is a two-stage security vetting benchmark for open agentic skill ecosystems, evaluating detection of malicious skills via semantic analysis and runtime verification in a sandbox. The benchmark is built from confirmed malicious skills in the OpenClaw ecosystem.
- Released
- 2026-05-30
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Open agent platforms face supply-chain risks from malicious skills, but existing defenses lack a standardized evaluation. This benchmark addresses the gap in measuring both detection and runtime verification, offering practical value for improving security in agent ecosystems.
Motivation
Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.