MIT CSAILSwinburneUCSDFeb 18, 2026arXiv:2602.16313

MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

Zexue He, Zexue He, Yu Wang, Churan Zhi, Churan Zhi, Yuanzhe Hu, Yuanzhe Hu, Tzu-Ping Chen, Tzu-Ping Chen, Lang Yin, Lang Yin, Ze Chen, Tong Arthur Wu, Jiaxin Pei, Siru Ouyang, Julian McAuley, Jiaxin Pei, Yejin Choi, A. Pentland, Yejin Choi, Alex Pentland

AI Summary

The paper introduces MemoryArena, a benchmark for evaluating agent memory in multi-session tasks where memorization and action are interdependent. MemoryArena comprises human-crafted agentic tasks with interdependent subtasks across domains like web navigation and formal reasoning, requiring agents to distill experiences into memory and use it to guide future actions. Experiments show that agents performing well on existing long-context memory benchmarks like LoCoMo struggle in MemoryArena, highlighting a gap in current memory evaluation methods.

Key Contribution

Agents that ace long-context recall can still bomb when they need to use that memory to actually *do* something, revealing a critical flaw in how we currently evaluate memory in AI.

Abstract

Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversations or text but fails to capture how memory is used to guide future decisions. Another class focuses on agents acting in single-session tasks without the need for long-term memory. However, in realistic settings, memorization and action are tightly coupled: agents acquire memory while interacting with the environment, and subsequently rely on that memory to solve future tasks. To capture this setting, we introduce MemoryArena, a unified evaluation gym for benchmarking agent memory in multi-session Memory-Agent-Environment loops. The benchmark consists of human-crafted agentic tasks with explicitly interdependent subtasks, where agents must learn from earlier actions and feedback by distilling experiences into memory, and subsequently use that memory to guide later actions to solve the overall task. MemoryArena supports evaluation across web navigation, preference-constrained planning, progressive information search, and sequential formal reasoning, and reveals that agents with near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly in our agentic setting, exposing a gap in current evaluations for agents with memory.

Eval Frameworks & Benchmarks Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References32

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

Related Papers