Search papers, labs, and topics across Lattice.
This paper introduces a novel recurrent Transformer architecture called Maglev, which utilizes fixed-size memory to enhance the efficiency of sliding-window attention while maintaining parallelization during training. By employing a two-model system鈥攚here a prefiller model generates expressive memory targets and a decoder model utilizes sliding-window attention with recurrent key/value injection鈥擬aglev achieves significant improvements in validation loss and performance on downstream pretraining tasks. The approach also benefits from parameter sharing between the two models, resulting in reduced memory usage without sacrificing performance gains.
Maglev achieves superior validation loss and performance benchmarks by cleverly combining full attention with sliding-window mechanisms, all while conserving memory.
We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. \ours{} consists of two coupled models: a prefiller $Q$, which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for $Q$, as this yields stronger performance. The essential requirement is that $Q$ be more expressive than $P$, with access to the full history.} to produce memory targets $m'_t$, and a decoder $P$, which uses only sliding-window attention and recurrent K/V injection to produce decoder memories $m_t$ for next-token prediction. We train \ours{} with a memory consistency loss that aligns $m_t$ with $m'_t$, allowing inference to use $P$ alone. Empirically, \ours{} improves validation loss and downstream pretraining benchmarks over sliding-window and latent recurrent transformer baselines. Moreover, sharing parameters between $P$ and $Q$ reduces parameter memory while preserving most of the gains.