Jun 11, 2026arXiv:2606.13681

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Jundong Xu, Qingchuan Li, Qingchuan Li, Jiaying Wu, Jiaying Wu, Yihuai Lan, Yihuai Lan, S. Li, Shuyue Stella Li, Huichi Zhou, Huichi Zhou, Bowen Jiang, Bowen Jiang, Lei Wang, Lei Wang, Jun Wang, Jun Wang, A. Luu, Anh Tuan Luu, Caiming Xiong, Caiming Xiong, Hae Won Park, Hae Won Park, Bryan Hooi, Bryan Hooi, Zhiyuan Hu, Zhiyuan Hu

AI Summary

This paper introduces EvoArena, a benchmark suite designed to evaluate large language model (LLM) agents in dynamic environments, addressing the limitations of static evaluations. The authors propose EvoMem, a memory paradigm that tracks memory evolution through structured update histories, enabling agents to adapt their knowledge and reasoning to changing conditions. Experimental results reveal that while current agents perform poorly on EvoArena, EvoMem enhances their performance by an average of 1.5%, demonstrating significant improvements in both evolving tasks and standard benchmarks.

Key Contribution

LLM agents struggle in dynamic environments, but EvoMem boosts their performance by capturing the evolution of memory, leading to better adaptability.

Abstract

Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deployment is inherently dynamic, requiring agents to continually align their knowledge, skills, and behavior with changing environments and updated task conditions. To address this gap, we introduce EvoArena, a benchmark suite that models environment changes as sequences of progressive updates across terminal, software, and social domains. We further propose EvoMem, a patch-based memory paradigm that records memory evolution as structured update histories, enabling agents to reason about environmental evolution through changes in their memory. Experiments show that current agents struggle on EvoArena, achieving an average accuracy of 39.6% across evolving terminal, software, and social-preference domains. EvoMem consistently improves performance, yielding an average gain of 1.5% on EvoArena and also improving standard benchmarks such as GAIA and LoCoMo by 6.1% and 4.8%. Beyond individual tasks, EvoMem further improves chain-level accuracy by 3.7% on EvoArena, where success requires completing a consecutive sequence of related evolutionary subtasks. Mechanistic analysis shows that EvoMem improves evidence capture in the memory, indicating better preservation of complete evolving environment states. Our results highlight the importance of modeling evolution in both evaluation and memory for reliable agent deployment.

Eval Frameworks & Benchmarks Tool Use & Agents World Models & Planning

Citation Metrics

Citations0

Influential citations0

References27

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Related Papers