Search papers, labs, and topics across Lattice.
This paper introduces OctoLong, a context engineering pipeline that leverages an AST parser, language server backend, and package manager to curate extensive, dependency-rich code contexts for training long-context language models. By mid-training OctoLong-Instruct on a mixture of 50 billion tokens, including 6.2 billion tokens from OctoLong, the authors demonstrate significant improvements in long-range retrieval and state tracking, outperforming 18 state-of-the-art long-context models. The results indicate that replacing just 12% of traditional context data with OctoLong contexts can enhance both long-term code understanding and short-context coding tasks.
Replacing just 12% of traditional training data with OctoLong's curated code contexts leads to substantial improvements in long-range retrieval and state tracking for language models.
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies. In this work, we introduce OctoLong, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of dependency-rich code contexts of millions of tokens in length. We then train OctoLong-Instruct, a suite of capable long-context open LMs, derived from base models ranging in size from 600M to 14B parameters, via context-extension mid-training on a ~50B-token mixture containing ~6.2B tokens of OctoLong code contexts, followed by ~10B tokens of instruction tuning. Our training ablations and experimental evaluations against 18 state-of-the-art open-weight long-context LMs show that supplanting just 12% of traditional context-extension corpora with OctoLong data yields substantial gains in long-range retrieval, long-term state tracking, repository-level code understanding, and downstream agentic tasks, while also enhancing API usage in short-context coding scenarios.