Search papers, labs, and topics across Lattice.
LOCKS introduces a novel approach to managing key-value (KV) caches in large language models by providing each page with its own compact spectral summary, significantly reducing the cache size needed for long-context decoding. This method allows for efficient reconstruction of within-page logits and selective attention to only the most relevant pages, achieving performance close to full KV caches while drastically reducing computational overhead. The results demonstrate that LOCKS can maintain high-quality outputs on long-document question answering and reasoning tasks while halving per-token decode latency, making it a practical solution for scaling long-context applications.
LOCKS achieves near-full KV performance with only 2% of the tokens attended, revolutionizing long-context decoding efficiency.
Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directions that a page's own compact basis retains. LOCKS gives every page its own spectral summary (resident, about a tenth the cache's size), reconstructs within-page logits, estimates each page's attention mass by log-sum-exp, and attends only the top pages; selection itself reads no candidate keys or values. Selecting on this summary alone stays within about a point of the full cache on long-document QA (LongBench-v1), tracks the read-every-key oracle on retrieval-dense RULER down to the smallest budgets, and shows its largest margins on long-form reasoning (AIME26, MATH-500), where baseline selectors collapse. At its shipped $2048$-token budget LOCKS matches FullKV aggregate quality at $100$K$+$ context while attending about $2\%$ of the tokens, and halves per-token decode latency ($2.0\times$ at $1$M tokens) against dense attention. LOCKS ships as a drop-in plugin for unmodified vLLM, with batched decode running in full CUDA graphs.