Search papers, labs, and topics across Lattice.
This paper introduces Skill-α, a reinforcement learning method that addresses the challenges of skill generation by framing it as a sequential editing process, allowing for the progressive construction of agent skills. By employing a novel rollback reward mechanism, Skill-α evaluates the impact of each edit on downstream task performance, ensuring that generated skills are both relevant and effective. Experimental results demonstrate that Skill-α outperforms traditional heuristic and pipeline-based methods, achieving significant improvements in success rates on benchmark tasks like CL-Bench and tau2-bench.
Skill-α outperforms traditional skill generation methods by leveraging a novel rollback reward mechanism, leading to significant improvements in agent performance on downstream tasks.
Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. In contrast, learning-based approaches offer a more unified way to model skill generation across heterogeneous sources. However, learning-based skill generation remains challenging because skills lack a natural supervision signal based on relevance or correctness; their value can largely be determined only by whether they improve the behavior of the agent on downstream tasks. To address this challenge, we propose Skill-α, a reinforcement learning method for progressively generating high-quality agent skills. Specifically, we formulate skill generation as a sequential editing process that decomposes skill construction into individually evaluable edits, and introduce a novel rollback reward that evaluates each edit by comparing downstream execution under the original and edited skills on an anchored query. Extensive experiments show that Skill-α generates more effective skills than methods based on heuristics or pipelines in both document-to-skill and experience-to-skill settings. Under the main GPT-4o worker, Skill-α improves average downstream success rates over the strongest skill-generation baseline by 3.3 points on CL-Bench and 6.7 points on tau2-bench. Further ablations validate the importance of rollback reward and progressive generation.