Search papers, labs, and topics across Lattice.
This paper introduces AFTER, a benchmark comprising 382 enterprise tasks across six professional roles to assess the effectiveness of procedural memory in LLM agents. The study reveals that procedural memory significantly enhances performance in industrial workflows, with a single refinement round yielding a 3.7-6.7 point improvement and multi-model execution traces achieving a 73.1% cross-model test accuracy. Notably, the research identifies a dichotomy in skill generalization, where some skills transfer effectively across tasks and models while others become overly specialized, offering critical insights for the deployment of procedural memory systems.
Procedural memory boosts LLM performance in workplace tasks, with some skills achieving over 73% accuracy across different models while others falter under transfer.
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones. The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization. Experiments show that procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points, while skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources. We further find that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer. These results provide practical guidance for building, evaluating, and deploying procedural memory systems in production agent platforms.