SkillHarm
Benchmark of skill-based attacks across the skill-use lifecycle, with 879 attack samples across 71 skills, evaluating Fixed-Payload Poisoning and Self-Mutating Poisoning scenarios across 12 risk types, with attack success rate as primary metric.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing studies evaluate poisoned skills within single task executions and use ad-hoc risk lists. This benchmark systematically covers lifecycle-aware attacks and provides a taxonomy and construction pipeline for reproducible evaluation.
Motivation
Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party skills a vulnerable attack surface.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.