Search papers, labs, and topics across Lattice.
Investigating the epistemic neutrality of generative models as political intermediaries, the authors perform activation steering on Llama 3.1 8B by leveraging its 2024 knowledge cutoff against subsequent political realignments as a natural experiment. Mechanistic probing reveals that alignment training merely masks a linearly encoded partisan direction in latent space rather than removing it. Consequently, the model presents temporally contingent political frames as objective facts, lacking the architectural capacity to separate empirical knowledge from shifting political consensus.
Alignment training doesn't erase political bias鈥攊t merely conceals a measurable latent direction that leads models to hardcode transient partisan consensus as timeless objective fact.
Large language models (LLMs) are rapidly becoming an interface between citizens and political information. They are often regarded as "a better Google." While this analogy might work for some instances, it is unintuitively problematic for democratic politics. A search engine retrieves human-authored documents, while a language model generates novel text that necessarily embeds invisible framing decisions. Because conveying knowledge involves framing, a system that generates answers cannot serve as a neutral conduit to "all human knowledge." Instead, these systems are becoming a new kind of political intermediary. Mechanistic evidence shows that partisan identity is encoded as a locatable geometric direction inside the Llama 3.1 8B model, and that alignment training masks rather than removes this structure. Building on that evidence, we present steering experiments that exploit a model's training cutoff in 2024. This cutpoint auspiciously falls just before a dramatic realignment in American politics marked by the second Trump administration and the MAHA transformation of health politics, providing us with a natural experiment. We find that the model presents temporally contingent partisan alignments as knowledge, with no mechanism for distinguishing fact from opinion. This reality moves the information environment beyond the echo chamber toward an epistemic monoculture where language models, purporting to summarize "all human knowledge" are, in actuality, simply magnifying the cultural and partisan divides inherent in their training data.