Search papers, labs, and topics across Lattice.
This paper introduces RING (Retrieval-Internalized Generation), a novel framework that integrates large-scale external knowledge into a Mixture-of-Memory Experts architecture while eliminating the need for external retrieval systems. By employing a three-stage training process that includes continued pre-training, supervised fine-tuning, and reinforcement learning, RING learns to optimize its internal memory retrieval policy directly from task signals. The results demonstrate that RING achieves comparable or superior accuracy and efficiency to traditional retrieval-augmented generation methods, while also addressing the latency and engineering challenges associated with external retrievers.
RING achieves superior knowledge integration without the latency of external retrieval, redefining efficiency in large-scale language models.
Retrieval-augmented generation (RAG) improves factuality but adds latency and engineering overhead at serving time. We propose RING (Retrieval-Internalized Generation), a holistic paradigm spanning both architecture and training that injects large-scale external knowledge into a \textit{Mixture-of-Memory Experts} and learns parametric search over this internal memory via reinforcement learning, removing the external retriever entirely. Training proceeds in three stages: continued pre-training injects new corpora into a Knowledge Expert via our novel \textit{Dual Causal Attention}; supervised fine-tuning teaches a ``search-then-answer'' pattern; and reinforcement learning with hierarchical rewards optimizes the routing-and-search policy over the parametric memory. Unlike prior parametric injection methods that pair internal memory with a fixed or rule-based retriever, RING {learns} its retrieval policy directly from task signals. We further frame RING theoretically as a search-free approximation to the classical RAG objective. To evaluate large-scale injection of genuinely {new} knowledge without test-time leakage, we further construct News-2025, a benchmark built from news strictly post-dating the base LLM's pretraining cutoff. RING matches or surpasses both search-based RAG and parametric injection baselines in accuracy and efficiency.