Search papers, labs, and topics across Lattice.
This paper introduces Zero-Mem, a novel approach to memory operations in LLM agents that eliminates the need for intermediate LLM calls during memory access, thereby reducing token consumption and processing time. By utilizing zero-token memory operations, Zero-Mem organizes interaction traces into an entity-context graph and a temporal hierarchy, allowing for efficient retrieval and contextual grounding without generating additional tokens. The method achieves a 57.6% reduction in memory-operation time costs compared to the fastest baseline while maintaining competitive performance on long-memory and long-context question-answering benchmarks.
Zero-Mem slashes memory operation time costs by over 57% while eliminating unnecessary LLM calls, revolutionizing how LLM agents manage memory.
LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original evidence. We ask whether structured memory access requires generation at all. Zero-Mem introduces \emph{zero-token memory operations}: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity--context graph exposes connections across interactions, while a temporal hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, retrieves from both, and follows their structure to recover supporting relations or surrounding context. Deterministic calibration first discards conflicting evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory and long-context question-answering benchmarks, Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations. With the same final-QA reader and context budget, it reduces memory-operation time cost by 57.6\% relative to the fastest compared baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, the results show that structured agent memory need not generate an intermediate representation of the past. After peer review, the code and implementation details will be available at \textcolor{blue}{https://github.com/TheMoon0815/Zero-mem}.