AI BENCHMARK PROFILE
WebDev-Skills-Bench
A benchmark for evaluating agent skills in web development, comparing matched conditions with length-matched controls and ablation studies on 31 skills across 50 projects.
- Released
- 2026-08-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Raises the standard for agent-skill evaluation by incorporating length-matched controls and per-model audits, revealing that skill efficacy varies by model and deployment.
Motivation
Agent Skills are reusable procedural modules that are increasingly injected into coding-agent sessions to encode framework conventions, anti-patterns, and reusable tools.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.