Search papers, labs, and topics across Lattice.
This study introduces WebDev-Skills-Bench, a benchmark designed to evaluate the effectiveness of Agent Skills in web development by analyzing their impact on task performance across various models. The empirical analysis of 31 public WebDev Skills on 50 projects reveals that Skill injection often leads to decreased task completion rates and increased token costs, with only a minority of Skill-project pairs showing positive outcomes. The findings highlight significant variability in Skill effectiveness, suggesting that the decision to inject Skills should be tailored to specific model and project contexts rather than treated as a universally beneficial practice.
Injecting Agent Skills in web development can reduce task performance by up to 4.2%, challenging the assumption that more Skills always lead to better outcomes.
Agent Skills are reusable procedural modules that are increasingly injected into coding-agent sessions to encode framework conventions, anti-patterns, and reusable tools. However, because each injected Skill expands the prompt of every query, an effective Skill benchmark must determine not only whether an agent can solve a task, but whether the Skill should have been injected at all. We introduce WebDev-Skills-Bench and use it for a controlled empirical study of 31 public WebDev Skills on 50 Web-Bench projects and 1,000 ordered tasks. The benchmark compares four matched conditions, including a length-matched irrelevant control and leave-one-out component ablations. To isolate Skill effects from prompt-length artifacts, we place only SKILL.md in the prompt while mounting auxiliary files into the agent workspace. Across four models, target Skill injection reduces mean Pass@2 by 1.3% to 4.2%, lowers task completion depth, and increases token cost by 72% to 394%, with gains in only 17% to 36% of Skill-project pairs. Length-matched controls reveal two failure modes: some models are length-distracted, where an equally long irrelevant Skill reproduces most of the loss, while others are content-misled, where prompt length is neutral but Skill content still lowers Pass@2 by 1.1% to 1.4%. Further analysis shows that losses concentrate on easy early tasks, Skill rankings transfer weakly across models, and anti-pattern rules outperform example-heavy content within helpful Skills. These findings recast a matched Skill as a hypothesis about a particular Skill-project-model triple rather than a portable asset, reframing injection as a per-deployment routing decision and making length-matched controls and per-model audits a minimum standard for Agent-Skill evaluation.