machine learning

Speculate with Memory: Lossless Acceleration for LLM Agents

arXiv:2607.12236

summary

The paper proposes adding online memory systems to speculative execution for large language model agents, enabling the speculator to learn from past trajectories and improve prediction accuracy without extra wall‑clock time.

Abstract

Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle. However, existing speculators are stateless and discard all information between tasks, preventing prediction quality from improving with experience. We equip the speculator with three online memory systems that learn from past agent trajectories: a contrastive transition table tracking action-sequence statistics, an episodic memory retrieving contextually similar segments, and a confusion tracker suppressing recurring errors. We evaluate this approach on six benchmarks spanning three speculation types: action prediction, observation prediction, and chained prediction. Memory-augmented speculation yields a 19--39\% relative accuracy improvement on action prediction and up to a increase on observation prediction tasks with repetitive action spaces. These gains grow continuously as memory accumulates and generalize across speculator models of varying cost. All speculation is lossless because it runs during idle time at zero added wall-clock cost, and the actor's trajectory is identical to non-speculative execution.

Topics & keywords

#speculative execution#large language models#memory augmentation#agent planning#online learningcontrastive transition tableepisodic memoryconfusion trackeraction predictionobservation predictionlossless acceleration
Speculate with Memory: Lossless Acceleration for LLM Agents · wovepaper