Search papers, labs, and topics across Lattice.
This paper emphasizes the importance of transitioning from short-term evaluations of language model interactions to longitudinal assessments that capture the cognitive, developmental, and socio-affective changes in users over time. By integrating social science measurement techniques with NLP methodologies, the authors aim to identify and mitigate long-term risks associated with human-AI interactions. The key finding is that understanding these behavioral shifts can enable proactive detection of issues, steering model development towards beneficial outcomes for users.
Long-term interactions with language models could lead to significant cognitive and emotional changes in users that short-term evaluations miss entirely.
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface in short-term interactions, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.