Benchmark Radar
AI BENCHMARK PROFILE

OpenSkillRisk

General AISafety & Trustworthiness

OpenSkillRisk evaluates LLM-based agents on their ability to recognize and avoid safety risks when using third-party skills, with 263 risky skills in seven threat categories and sandboxed execution.

Released
2026-07-22
Readiness
Paper only
Primary field
General AI

Why it matters

Third-party skills can introduce latent security vulnerabilities. A dedicated safety benchmark helps assess agent risk reasoning and execution control in open-world scenarios.

Motivation

LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.