Benchmark Radar
AI BENCHMARK PROFILE

MaliciousSkillBench

General AIMultimodal PerceptionProtectSkills

Benchmark for detecting malicious agent skills, comprising 9,740 skills (7,505 malicious, 2,235 benign) and evaluating detectors under random, structural-disjoint, and source-disjoint splits.

Released
2026-08-20
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the need for robust malicious skill detection by providing a consolidated cross-source dataset and evaluation protocols that measure both detection and false-positive rates.

Motivation

Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.