Search papers, labs, and topics across Lattice.
This paper introduces FinPerMA, a novel benchmark designed to evaluate the personalized memory capabilities of large language model (LLM) agents in high-stakes contexts like financial advising. By utilizing a combination of deterministic impact rules and controlled LLM narration, the benchmark assesses how well LLMs can integrate significant events into an evolving user model over time. The findings reveal that current LLMs struggle with maintaining personalized memory, achieving only about 39% accuracy on multiple-choice questions, indicating a critical gap in their ability to adapt to user preferences after significant events.
Despite advances in LLMs, they fail to effectively integrate user preferences over time, with accuracy rates stagnating around 39% even in ideal conditions.
Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.