Benchmark Radar
AI BENCHMARK PROFILE

SkillVetBench

General AIMultimodal Perception

SkillVetBench is a two-stage security vetting benchmark for open agentic skill ecosystems, evaluating detection of malicious skills via semantic analysis and runtime verification in a sandbox. The benchmark is built from confirmed malicious skills in the OpenClaw ecosystem.

Released
2026-05-30
Readiness
Paper only
Primary field
General AI

Why it matters

Open agent platforms face supply-chain risks from malicious skills, but existing defenses lack a standardized evaluation. This benchmark addresses the gap in measuring both detection and runtime verification, offering practical value for improving security in agent ecosystems.

Motivation

Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.