Search papers, labs, and topics across Lattice.
This paper introduces a memory-augmented speculative execution framework for LLM agents that utilizes three online memory systems to enhance prediction quality by retaining information across tasks. By implementing a contrastive transition table, episodic memory, and a confusion tracker, the approach achieves significant improvements in action and observation prediction accuracy on various benchmarks. The results demonstrate a 19-39% relative accuracy increase in action prediction and up to a 2.5x enhancement in observation prediction tasks, all while maintaining lossless execution during idle times.
Memory-augmented speculation boosts LLM prediction accuracy by up to 39% without incurring any additional execution time.
Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle. However, existing speculators are stateless and discard all information between tasks, preventing prediction quality from improving with experience. We equip the speculator with three online memory systems that learn from past agent trajectories: a contrastive transition table tracking action-sequence statistics, an episodic memory retrieving contextually similar segments, and a confusion tracker suppressing recurring errors. We evaluate this approach on six benchmarks spanning three speculation types: action prediction, observation prediction, and chained prediction. Memory-augmented speculation yields a 19--39\% relative accuracy improvement on action prediction and up to a $2.5\times$ increase on observation prediction tasks with repetitive action spaces. These gains grow continuously as memory accumulates and generalize across speculator models of varying cost. All speculation is lossless because it runs during idle time at zero added wall-clock cost, and the actor's trajectory is identical to non-speculative execution.