Benchmark Radar
AI BENCHMARK PROFILE

SkillResolve-Bench

General AIKnowledge & ReasoningSearch & Retrieval

SkillResolve-Bench 1.0 evaluates agent skill retrieval under same-capability ambiguity, pairing helpful skills with risky siblings. It includes 661 pairs, a 7,982-candidate pool, disjoint splits, and reports Recall@K and harmful sibling rate (HSR@K).

Released
2026-06-09
Readiness
Paper only
Primary field
General AI

Why it matters

Skill retrieval carries execution risk beyond relevance. This benchmark quantifies exposure to risky siblings, supporting development of retrievers that select safe representatives, reducing harmful failures in agent deployments.

Motivation

Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and execution assumptions to an agent.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.