Search papers, labs, and topics across Lattice.
This study investigates the origins of output homogeneity in language models (LMs), arguing that such convergence is primarily learned during the pretraining phase rather than solely during alignment. Through controlled instruction-tuning experiments, the authors demonstrate that while the alignment process amplifies this homogeneity, it does not initiate it, as convergence can be induced by prompting base models alone. These findings imply that the semantic convergence observed in LMs is an inherent characteristic of their training objectives, complicating efforts to address it post-alignment.
Semantic convergence in language models may be an inherent trait from pretraining, not just a byproduct of alignment, challenging conventional beliefs about output diversity.
The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only \emph{revealed} or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in base models. We find that instruct-like collapse can be induced through prompting alone, even without alignment. Taken together, our results suggest that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.