Benchmark Radar
AI BENCHMARK PROFILE

MalSkillBench

Transport & LogisticsKnowledge & Reasoning

MalSkillBench is a runtime-verified benchmark of malicious agent skills, with 3,944 malicious skills labeled along a taxonomy, measuring detection tool effectiveness.

Released
2026-06-05
Readiness
Paper only
Primary field
Transport & Logistics

Why it matters

Evaluates detection tools for hybrid code-prompt skills, potentially informing supply chain security, but lacks public artifacts for reuse.

Motivation

AI coding agents such as Claude Code and Gemini CLI increasingly extend themselves with third-party skills: markdown packages bundling natural-language instructions, executable scripts, and tool permissions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.