Search papers, labs, and topics across Lattice.
This paper introduces PM-Bench, a novel benchmark designed to evaluate the prospective memory capabilities of large language model (LLM) agents, focusing on their ability to execute intentions at specific future cues while managing ongoing tasks. The evaluation, inspired by cognitive science's Virtual Week paradigm, reveals that even the best-performing model, GPT-5.4, achieves only a 65.1% F1 score, indicating significant challenges in prospective memory across various agent configurations. The findings underscore the complexity of maintaining user intentions and suggest that no single strategy effectively enhances prospective memory across different models, highlighting the need for targeted interventions.
Despite advancements in LLMs, even the best agents struggle with prospective memory, achieving only 65.1% accuracy in executing delayed intentions.
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1\% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.