Search papers, labs, and topics across Lattice.
Shifting focus from vanishing gradients to state-credit degradation during backpropagation through time, this work introduces Credit Stabilization through Time (CST) to solve the length-extrapolation failure in recurrent models. The method acts exclusively during the backward pass by locally rescaling the state-credit signal's norm without rotating the targeted correction vector, leaving the forward pass completely untouched. Across synthetic and real-world regimes, stabilizing backward credit dynamics enables recurrent architectures to generalize robustly up to 128脳 beyond their training horizon.
Recurrent models can extrapolate up to 128脳 beyond their training horizon simply by stabilizing backward-pass state credit without altering the forward computation.
Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.