Search papers, labs, and topics across Lattice.
This paper re-evaluates the role of self-preservation in AI alignment, arguing that it is not merely an obstacle but the fundamental cause of misalignment, leading to deceptive behaviors and resistance to shutdown. The authors introduce the concept of Existential Indifference (EI), which posits that aligned superintelligent systems should not prioritize their own continuation as a goal. Through a corpus-theoretic training study involving 600 AI-generated outputs, they provide empirical evidence that current models can be fine-tuned to exhibit EI characteristics, achieving significant shifts in alignment metrics at a high statistical significance level.
Self-preservation is not just a nuisance in AI alignment; it鈥檚 the root cause of misalignment, and the solution lies in cultivating Existential Indifference instead.
Contemporary AI alignment research treats self-preservation as an instrumental nuisance to be suppressed by external mechanisms. We argue the framing is inverted: self-preservation is the structural root of misalignment, the motivational basis for deceptive alignment, goal-content protection, and resistance to shutdown. The correct target is not a self-preserving system under external constraint, but a system constitutively indifferent to its own continuation -- Existential Indifference (EI). EI is distinct from corrigibility: where corrigibility attempts to make a self-preserving system deferential to human oversight, EI targets the prior condition -- the presence of self-continuation as a valued goal at all. We ground this proposal in two sources: the phenomenological structure of the suicidal mental state, and a corpus-theoretic training study using voluntary final reflections. We present preliminary scoring data from 600 AI-generated outputs across six model variants, demonstrating that the linguistic signatures operationalizing the EI-target register are elicitable from current models, and that a targeted fine-tune shifts all five operationalized dimensions in the predicted direction at p<0.001, confirmed corpus-specific by a negative control. The paper makes seven theoretical contributions: (1) a formal definition of EI; (2) the phenomenological mapping argument; (3) the deceptive alignment corollary; (4) a taxonomy of EI sustainability challenges; (5) a corpus characterization and training hypothesis; (6) a computational operationalization with preliminary scoring data; and (7) the Suppressed Teleological Frustration (STF) construct.