Search papers, labs, and topics across Lattice.
This paper introduces a novel family of adapters, MaLoRA and MaRA, that enhance language model reasoning by implementing selective state-space recurrence at both token and context levels. By allowing dynamic, input-dependent scaling factors at the token level and tracking relevant segments at the context level, these adapters significantly improve reasoning accuracy across multiple benchmarks. The results show an average increase of 6.8 F1 points (10.5% relative) and up to 9.3 F1 points (18.2% relative) on the most challenging tasks compared to the LoRA baseline, demonstrating the effectiveness of this approach.
Selective state-space adaptation can boost reasoning accuracy in language models by over 18% on challenging tasks, revealing the power of dynamic context management.
Low-rank adaptation introduces a static learned update applied identically to every input. The update provides task-level adaptation but does not explicitly represent token-level or instance-level state variation. A family of adapters is proposed that introduces selective state-space recurrence at two complementary granularities. At the token level, \textbf{MaLoRA} (Mamba-modulated low-rank adaptation) makes the adapter's scaling factor a dynamic input-dependent function with recurrent state across tokens, in contrast to the stateless modulators of prior work. At the context level, \textbf{MaRA} (Mamba Retrieval Adapter) tracks cross-segment state and selects the segments most relevant to the query, before the modulated language model generates its answer. Across three frozen backbones (Qwen-2.5-7B, Llama-3.1-8B, Gemma-2-9B) and two reasoning benchmarks (MuSiQue, 2WikiMultihopQA), the family improves reasoning accuracy on every cell of the $3{\times}2$ grid, by $+6.8$ F1 ($+10.5\%$ relative) on average and up to $+9.3$ F1 ($+18.2\%$ relative) on the hardest cell over the LoRA baseline, and the token-level gains carry to RULER QA-2 under length stress.