Search papers, labs, and topics across Lattice.
This paper introduces Cooperative Parameter-subspace Evolution Strategy (CoPES), a novel method that enhances the efficiency of post-training for resource-constrained tool-using large language model (LLM) agents by decomposing the full parameter space into lower-dimensional subspaces. CoPES achieves 92% of the validation accuracy gain of gradient-based reinforcement learning (GRPO) while requiring less than one-eighth of the GPU memory, significantly outperforming standard evolution strategies and LoRA-based GRPO across multiple benchmarks. The findings indicate that CoPES not only reduces training time but also improves the trade-off between memory usage and performance in agentic LLM post-training.
CoPES recovers 92% of the performance gains of traditional methods while slashing GPU memory requirements to less than one-eighth.
Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eventually match the performance of gradient-based reinforcement learning (RL). However, resource-constrained settings typically offer only a few GPUs, so the high GPU-hour requirements of ES translate into prohibitively long training times. To address this, we introduce Cooperative Parameter-subspace Evolution Strategy (CoPES), a cooperative coevolutionary method that decomposes the full parameter space into lower-dimensional subspaces and searches over them cooperatively to improve optimization efficiency. We post-train a Qwen3.5-4B tool-using agent for the math task and evaluate it on five benchmarks of varying difficulty. Under the GPU-hour budget of full-parameter GRPO's best validation checkpoint, CoPES recovers 92% of GRPO's validation-accuracy gain, versus 67% for standard ES, while its theoretical GPU memory requirement is less than one-eighth that of full-parameter GRPO. It consistently outperforms standard ES and LoRA-based GRPO on all evaluated pass@k metrics across the five benchmarks. Additional experiments further show the advantage of CoPES on the question-answering task. These results demonstrate an improved trade-off between memory requirements and training time for agentic LLM post-training under resource constraints. The code is open-sourced in https://github.com/MetaronWang/CoPES