Search papers, labs, and topics across Lattice.
This paper introduces Latent-LoRA, a novel approach to continual learning that leverages pooled token embeddings from a frozen LLM to facilitate task-agnostic adapter selection without gradient-based training. By employing a Gaussian mixture model for routing and constraining task parameters to the principal subspace of pretrained weights, the method achieves significant reductions in parameter usage while maintaining performance. Experiments demonstrate that Latent-LoRA outperforms existing methods with near-zero forgetting across multiple model scales and benchmarks, highlighting its effectiveness in mitigating catastrophic forgetting in large language models.
Task-agnostic adapter selection without trainable parameters leads to state-of-the-art continual learning performance with minimal forgetting.
Large language models generalize well to individual tasks but lack an inherent mechanism for learning them sequentially, leading to catastrophic forgetting. To mitigate this, LoRA-based continual learning methods allocate a separate low-rank adapter per task, yet existing approaches either require task identity at inference or sum all adapters indiscriminately, letting irrelevant branches distort the output. Recent gating-based solutions route inputs to the correct adapter but introduce trainable parameters that themselves need protection against forgetting. In this work, we observe that pooled token embeddings from a frozen LLM embedding layer already separate task distributions throughout the learning sequence. A Gaussian mixture model fitted on these embeddings, without any gradient-based training, is sufficient for task-agnostic adapter selection at test time. This eliminates the need for a learned gating module. On the adapter side, constraining each task's parameters to the principal subspace of the pretrained weights via SVD yields a compact latent-space parameterization. Within this subspace, orthogonal regularization directly controls inter-task interference. The resulting system, Latent-LoRA, is replay-free, requires no trainable routing component, and uses substantially fewer parameters per task. Experiments across five model scales and two established continual learning benchmarks show state-of-the-art performance with near-zero forgetting.