AI BENCHMARK PROFILE
OpenSkillRisk
OpenSkillRisk evaluates LLM-based agents on their ability to recognize and avoid safety risks when using third-party skills, with 263 risky skills in seven threat categories and sandboxed execution.
- Released
- 2026-07-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Third-party skills can introduce latent security vulnerabilities. A dedicated safety benchmark helps assess agent risk reasoning and execution control in open-world scenarios.
Motivation
LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.