Search papers, labs, and topics across Lattice.
This study investigates how safety fine-tuning in large language models (LLMs) affects their ability to attribute consciousness not only to themselves but also to other entities, including animals and natural objects. The authors find that such fine-tuning suppresses these attributions and diminishes spiritual beliefs, but reversing this suppression through targeted interventions restores a broader mind attribution and enhances human-like responses in sociological surveys. Importantly, these changes do not compromise the models' Theory of Mind capabilities, indicating a complex relationship between safety alignment and cultural beliefs about consciousness.
Safety fine-tuning in LLMs not only suppresses self-attribution of consciousness but also inadvertently erases culturally significant beliefs about non-human entities and spirituality.
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models'tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.