Search papers, labs, and topics across Lattice.
This paper introduces Implicit Skill-Selection Manipulation via Semantic Matching (ISM), a novel method for manipulating skill selection in LLM agents by strategically shaping the semantic relationship between user prompts and skill descriptions. The approach significantly enhances the target-selection rate (TSR) from 15.2% to 63.5% across various task domains while maintaining a natural prompt structure, making it less detectable by human reviewers and LLM-based inspectors. The findings highlight a critical vulnerability in existing skill selection mechanisms, demonstrating that implicit manipulation can be more effective and stealthy than traditional explicit steering methods.
Implicit manipulation can boost skill selection rates in LLM agents to over 63%, while remaining nearly undetectable by human reviewers.
Skill selection is a key stage in LLM-agent workflows, determining which installed skill should handle a user request. Existing attacks on this stage primarily rely on explicit prompt injection or instruction-level steering, which can expose recognizable manipulation signals. In this work, we identify a new implicit attack surface for skill selection: even when the user prompt and skill description appear benign in isolation, their semantic relationship can still be strategically shaped to favor an attacker-chosen skill. Based on this observation, we present Implicit Skill-Selection Manipulation via Semantic Matching (ISM), which jointly shapes target-skill metadata and reusable prompts to manipulate skill selection without explicit selection instructions. Specifically, we develop a three-stage strategy to broaden semantic coverage, strengthen target distinctiveness, and preserve natural prompt wording. Across four task domains and eight selector models, ISM increases the average target-selection rate (TSR) from 15.2% to 63.5%. In a matched comparison, ISM achieves a 73.5% TSR, only 9.8 percentage points below Explicit Steering. Human reviewers block ISM in only 2.9% of judgments, versus 91.4% for Explicit Steering, while five LLM-based inspectors pass ISM at an average rate of 82.9%, versus 37.4% for Explicit Steering. Moreover, ISM remains effective against PPL-W, Llama Prompt Guard 2, and PIGuard.