Search papers, labs, and topics across Lattice.
FlowEvo introduces a novel framework where workflows and executable skills co-evolve in real-time during inference, allowing agents to adaptively construct and refine their capabilities. By compiling successful workflows into a persistent skill bank and suppressing those that lead to negative transfer, FlowEvo enhances the agent's performance without requiring extensive training. The method significantly outperforms existing baselines, achieving an 85.6% accuracy on ALFWorld, which is 26.4 points higher than the next best model while utilizing only one-third of the tokens.
FlowEvo enables agents to continuously evolve their skills and workflows on-the-fly, leading to unprecedented performance gains in complex task execution.
Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution. Existing skill libraries provide reusable executable routines, but are typically assembled offline and do not grow from the agent's own workflows. We introduce FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time. FlowEvo compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows. It also tracks each skill's downstream utility and suppresses skills that cause negative transfer. Using a shared GPT-4o-mini backbone, FlowEvo achieves the highest accuracy among 8 baselines on the full standard splits of ALFWorld, HumanEval, MBPP, GSM8K, and MATH-500. On ALFWorld, it reaches 85.6%, 26.4 points above the strongest baseline, while using roughly one third as many tokens. Across 10 base models spanning 7B to 671B parameters, FlowEvo outperforms ExpeL in 49 of 50 model-dataset comparisons. Code is available at https://github.com/DEFENSE-SEU/FlowEvo.