Search papers, labs, and topics across Lattice.
This paper introduces SkillGate, a novel approach to training long-horizon agents to select skills effectively during execution by addressing the issue of selector credit starvation. By partitioning credit into two distinct channels鈥攐utcome credit for execution tokens and action-local advantage for skill-naming tokens鈥擲killGate enables agents to make more informed skill selections without being misled by prior failures. The results demonstrate a significant improvement in trial success rates, increasing from 40.8% to 53.2% on five benchmarks, while reducing exposure to misleading candidates and skill usage.
SkillGate achieves a 30% boost in trial success for long-horizon agents by fundamentally rethinking how skill selection is rewarded during execution.
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence-level advantage, the few tokens that name the chosen skill carry a vanishing share of the loss, and the credit they inherit is increasingly wrong-signed as trajectories lengthen. A correct choice is punished whenever the execution after it fails, even though the choice itself is among the most valuable decisions in the trajectory. Auditing a completed run's own training artifacts confirms all three properties, each worsening monotonically with horizon. SkillGate removes the failure by construction: it partitions the token support into two disjoint credit channels, outcome credit reaching only execution tokens, and a separate action-local advantage reaching exactly the skill-naming tokens, positive only when a trajectory's single read is the correct one. On five agentic benchmarks under a 16-candidate slate, SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.