Search papers, labs, and topics across Lattice.
This paper introduces SPEE (Self-Progressive Experience Evolution), a novel framework that integrates experience distillation with policy optimization to enable self-improvement in large language models (LLMs). By employing a two-stage process of explicit experience evolution and implicit policy optimization, SPEE effectively transforms transient interaction experiences into persistent model capabilities, addressing the limitations of existing self-improvement paradigms. Experimental results across five mathematical reasoning benchmarks show that SPEE significantly outperforms both test-time and training-time self-evolution approaches, highlighting its effectiveness in enhancing model performance through experience accumulation.
Bridging the gap between transient experience and persistent capabilities, SPEE enables LLMs to self-improve by evolving their knowledge through a unique experience distillation process.
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training-time optimization methods can update model parameters but lack an explicit mechanism for accumulating transferable experience. Bridging these two paradigms requires a critical intermediate stage that remains underexplored, namely \emph{experience distillation}. To address this gap, we propose \textbf{SPEE} (\textbf{S}elf-\textbf{P}rogressive \textbf{E}xperience \textbf{E}volution), a unified post-training framework that sequentially performs explicit experience evolution followed by implicit policy optimization. During explicit experience evolution, SPEE reflects on trajectories collected from multiple interactions to extract, verify, and progressively evolve transferable experience, which is subsequently internalized into the policy through privilege-guided On-Policy Self-Distillation (OPSD). During implicit policy optimization, reward-driven reinforcement learning leverages these internalized priors to explore novel solution strategies. In the experience evolution stage, a continuously evolving global experience pool consolidates knowledge from both successful and failed trajectories, filters out low-utility experience, and mitigates post-hoc rationalization induced by individual trajectories. Experiments on five mathematical reasoning benchmarks demonstrate that SPEE consistently outperforms both test-time and training-time self-evolution baselines across three model scales. The source code is available at https://github.com/rrrsj/SPEE.