Search papers, labs, and topics across Lattice.
This paper develops a mathematical modeling framework grounded in social physics to explore the long-term implications of AI alignment in the context of evolving social norms. By combining analytical and simulation approaches, the authors investigate potential risks such as value lock-in and normative mode collapse that arise from static alignment strategies. The findings underscore the necessity for adaptive alignment mechanisms to ensure that AI systems remain responsive to changing human values and societal contexts.
Value lock-in from static AI alignment could jeopardize our ability to adapt to evolving social norms, risking a normative mode collapse in future AI interactions.
AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of social physics models as an epistemic bridge: enabling rapid, rigorous, and quantitatively-grounded hypothesis testing for sociotechnical foresight in general AI futures, and acting as a tractable precursor to more computationally expensive large-scale agentic evaluations.