Search papers, labs, and topics across Lattice.
This study investigates the role of auxiliary views鈥攔eformulations of knowledge鈥攊n the knowledge acquisition process of large language models (LLMs) during pre-training. Through controlled experiments, the authors demonstrate that repetition is essential for knowledge acquisition, and that paraphrasing enhances learning particularly at smaller batch sizes. Notably, they find that reallocating tokens from document repetition to auxiliary views improves factual recall, regardless of the teacher model's strength, highlighting the significance of data diversity in pre-training success.
Auxiliary views can enhance LLM learning efficiency, revealing that reallocation of training tokens can lead to better factual recall even when using weaker teacher models.
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.