Search papers, labs, and topics across Lattice.
This paper surveys the evolving role of memory in large language models (LLMs), establishing a systematic taxonomy that categorizes memory mechanisms based on representation, update dynamics, and persistence. By formalizing the processes of memory writing, routing, state transitions, and consolidation, the authors bridge the gap between computation-coupled and independently addressable memory architectures. The framework not only clarifies the fragmented landscape of memory strategies but also sets the stage for future advancements in scalable and adaptive language modeling.
A unified taxonomy reveals how diverse memory mechanisms in LLMs can be systematically understood and leveraged for future innovations.
Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.