Search papers, labs, and topics across Lattice.
This paper introduces Chained Recursive Language Models (Chained RLM), an innovative architecture that enhances long context reasoning in large language models by breaking down complex tasks into manageable sub-tasks. By utilizing a sequence of fresh reasoning roots that rely on compact summaries and task-specific artifacts, the model mitigates the propagation of errors that typically occur in single inference trajectories. The evaluation demonstrates that this method significantly improves accuracy in multi-hop reasoning tasks compared to traditional direct answering approaches, highlighting its effectiveness in managing context and intermediate states.
Chained RLMs can boost accuracy in multi-iteration reasoning tasks by effectively managing context and correcting errors through a novel inference architecture.
Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer. This becomes particularly difficult in tasks that require extraction, counting, ordering, or multi-hop reasoning, where an early mistake can propagate until the final response. In this work, we propose Chained Recursive Language Models (Chained RLM), an inference-time architecture, in which the same underlying model is called repeatedly as a sequence of fresh reasoning roots. Each root receives the original problem and context, but does not inherit the full conversational history. Instead, it receives a compact plain-text summary, a plain-text blackboard, and some durable task-specific artifacts written by predecessor roots. The motivation is to manage the context by chopping into partial tasks rather than one large inference response; in each staged computation, intermediate artifacts can be inspected, corrected, and extended by a later fresh inference by the same model. We describe the system model, handoff mechanism, artifact workspace, and evaluation protocol for this system. We study when fresh-context artifact continuation gives a measurable gain in accuracy over direct LLM answering even with recursive tool-calling.