Search papers, labs, and topics across Lattice.
This paper introduces RUMBA, a novel benchmark designed to assess long-term conversational memory in large language models (LLMs), particularly focusing on the Russian language. By providing a detailed taxonomy of memory-centric question types and a unified methodology that incorporates semantic types, session scope, and temporal reasoning, RUMBA addresses the limitations of existing English-centric benchmarks. Evaluation of contemporary memory systems reveals that RUMBA effectively diagnoses model behavior, highlighting both strengths and weaknesses in memory mechanisms across various contexts.
RUMBA reveals critical insights into how memory mechanisms in LLMs perform across long-term contexts, exposing significant gaps in current models' capabilities.
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions. RUMBA consists of timestamped user-assistant dialogues with QA pairs requiring retrieval, combination, and reasoning across sessions. While designed for Russian, we also provide an aligned English subset under the same methodology. We evaluate contemporary memory systems and long-context models, and show how RUMBA serves as a diagnostic tool to analyze model behavior across benchmark slices and identify strengths and failure modes of different memory mechanisms.