Benchmark Radar
AI BENCHMARK PROFILE

WebDev-Skills-Bench

General AIKnowledge & Reasoning

A benchmark for evaluating agent skills in web development, comparing matched conditions with length-matched controls and ablation studies on 31 skills across 50 projects.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

Raises the standard for agent-skill evaluation by incorporating length-matched controls and per-model audits, revealing that skill efficacy varies by model and deployment.

Motivation

Agent Skills are reusable procedural modules that are increasingly injected into coding-agent sessions to encode framework conventions, anti-patterns, and reusable tools.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.