Search papers, labs, and topics across Lattice.
This paper introduces ContextWeave, a longitudinal benchmark designed to evaluate the impact of memory on the performance of language agents in realistic, multi-task office workflows. By reconstructing workflows from 14 participants into 1,005 executable tasks, the study assesses how recalled experiences influence agent performance, revealing that a well-structured memory can significantly enhance both Workspace and Preference Scores. The findings indicate that while experience-rich memory improves workflow efficiency, it also poses risks of misleading recall, underscoring the need for advanced memory systems that balance retrieval relevance with execution reliability.
Experience-rich memory boosts agent performance in office workflows but can also lead to misleading recall, challenging traditional evaluation methods.
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.