Search papers, labs, and topics across Lattice.
This paper introduces VoiceMem, a novel memory architecture for conversational systems that integrates a dual-brain approach鈥攁n informational left brain and an emotional right brain鈥攁long with efficient streaming memory I/O mechanisms. The system significantly enhances accuracy, outperforming traditional memory systems by nearly 30 points in top-5 retrieval tasks, while also achieving state-of-the-art performance in emotional and personal interaction benchmarks. Additionally, VoiceMem operates in real-time with a retrieval time of 134 ms, ensuring low latency and cost-effectiveness for real-world applications.
VoiceMem's dual-brain architecture boosts conversational accuracy and emotional engagement, setting a new standard for real-time interaction in speech language models.
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional&Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time&Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.