Search papers, labs, and topics across Lattice.
This paper introduces Doc-to-Atom (Doc2Atom), a novel compositional memory framework that decomposes documents into semantically typed knowledge atoms, each represented by independent micro-LoRA adapters. By employing a lightweight query router to assemble relevant atoms for specific queries, the approach addresses issues of irrelevant-query interference and enhances scalability for long-document reasoning. Experimental results across six QA benchmarks show that Doc2Atom significantly outperforms existing methods while reducing memory costs associated with document internalization.
Compiling documents into semantically typed knowledge atoms allows for targeted, efficient retrieval that boosts performance on long-document reasoning tasks.
Long input sequences are central to document understanding and multi-step reasoning in Large Language Models, yet the quadratic cost of attention makes inference both memory-intensive and slow. Context distillation mitigates this by compressing contextual information into model parameters, and recent work such as Doc-to-LoRA amortizes context distillation into a single forward pass that generates one LoRA adapter per document. However, producing a single monolithic adapter for all queries leads to irrelevant-query interference, limited compositional recall, and poor scalability to long-document reasoning. To address these challenges, we propose Doc-to-Atom (Doc2Atom), a compositional parametric memory framework that decomposes each document into semantically typed knowledge atoms. Each atom is compiled into an independent micro-LoRA adapter and a provenance retrieval key. At inference time, a lightweight query router selects and assembles only the relevant atoms into a query-specific adapter, which is then injected into a frozen base model. The entire system is trained end-to-end through a multi-objective distillation framework. Experiments on six diverse QA benchmarks demonstrate that Doc2Atom outperforms Doc-to-LoRA baselines while reducing the memory cost of document internalization.