Search papers, labs, and topics across Lattice.
This paper introduces MemTrapBench, a novel benchmark designed to evaluate cognitive traps in large language models (LLMs) related to memory use, specifically focusing on how retrieved memories can distort reasoning and degrade task performance. The authors identify two key cognitive traps鈥擱easoning Fixation and Belief Distortion鈥攁nd demonstrate that all tested memory strategies underperform compared to a no-memory baseline, with significant performance drops exceeding 10%. To address these issues, they propose AdaptiveMem, an effective inference-time method that helps LLMs avoid cognitive traps while maintaining or enhancing performance on standard memory benchmarks.
Cognitive traps in LLM memory can lead to over 10% performance degradation, challenging the assumption that more memory always improves reasoning.
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.