Search papers, labs, and topics across Lattice.
This study investigates how large language models (LLMs) organize moral knowledge by training six independent linear probes corresponding to the categories of Moral Foundations Theory (MFT). The findings reveal that LLMs do not merely detect moral content but represent moral foundations in a complex geometric structure that spans multiple independent dimensions, indicating a shared component of integration specific to moral concepts. Notably, the model's representation of moral dilemmas captures the inherent tension within moral conflicts rather than providing pre-resolved judgments, highlighting the nuanced understanding of morality encoded in LLMs.
LLMs reveal a complex geometric structure of moral knowledge that captures the tension in moral dilemmas, rather than offering simplistic moral resolutions.
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.