Benchmark Radar
AI BENCHMARK PROFILE

Skill-Use

General AIKnowledge & Reasoning

Skill-Use evaluates skill use in agentic harnesses through 79 real skills and 177 executable tasks across nine domains, measuring trigger, compliance, and boundary adherence with a combined SU score.

Released
2026-08-05
Readiness
Paper only
Primary field
General AI

Why it matters

Agent evaluations often focus on task success, not whether agents can autonomously identify and apply relevant skills. Skill-Use isolates skill retrieval and usage under progressive disclosure, showing harness-dependent capabilities.

Motivation

Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.