Search papers, labs, and topics across Lattice.
This paper introduces SkillGate, a security gateway designed to detect malicious skill files in AI coding agents, addressing a significant vulnerability in the software supply chain. By employing a hybrid regex-prefilter combined with an LLM-judge pipeline, SkillGate effectively identifies and screens potentially harmful skill packages while minimizing runtime overhead and false positives. The results demonstrate that SkillGate achieves an F1 score of 0.817 and a false positive rate of just 1.13%, significantly outperforming existing detection tools while reducing LLM input tokens by 77%.
SkillGate reduces the risk of malicious skill files in AI coding agents by achieving an impressive F1 score of 0.817 while slashing LLM input requirements by 77%.
Software engineering teams now deploy AI coding agents (Cursor, Claude Code, GitHub Copilot) as first-class productivity tools, installing domain-specific skill files to tailor agent behavior to project APIs, framework conventions, and organizational workflows. These complex Markdown files are easily downloaded from public registries with a single npx skills add command and no real security screening, representing a novel supply-chain attack surface: a malicious skill file can silently reprogram agent behavior, exfiltrating credentials, injecting backdoors into generated code, or redirecting agent actions to attacker-controlled endpoints. The threat is not hypothetical: recent reports document hundreds of malicious skill packages in public registries, including organized campaigns that distributed credential-stealing infostealers via fake productivity skills. No systematic toolchain defense exists for this attack surface. We present SkillGate, a deployable security gateway that screens AI skill packages before coding agent installation. SkillGate uses a hybrid regex-prefilter + LLM-judge pipeline: safe-signal files bypass the LLM entirely (skip savings); flagged files have only their matched snippet windows sent to the judge, not the full content (snippet savings). We answer four research questions covering detection effectiveness, screening cost, runtime overhead, and false positive behavior on the SkillsBench benchmark against two existing tools. On SkillsBench (n=1,650, 9.1% malicious), SkillGate achieves F1=0.817, FPR=1.13% while reducing LLM input tokens by 77% vs. full-file screening, and outperforming existing tools by 5-6x on threshold-independent AUPRC (0.830 vs. 0.144/0.162).