AI BENCHMARK PROFILE
Skill-Use
Skill-Use evaluates skill use in agentic harnesses through 79 real skills and 177 executable tasks across nine domains, measuring trigger, compliance, and boundary adherence with a combined SU score.
- Released
- 2026-08-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Agent evaluations often focus on task success, not whether agents can autonomously identify and apply relevant skills. Skill-Use isolates skill retrieval and usage under progressive disclosure, showing harness-dependent capabilities.
Motivation
Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.