Search papers, labs, and topics across Lattice.
This study investigates the implicit personalization behavior of large language models (LLMs) in response to demographic cues, revealing a strong correlation (up to r=0.87) between internal activation signals and output changes in recommendations. By employing matched cued and neutral conversations across five LLMs, the authors demonstrate that the influence of demographic cues can be selectively suppressed by manipulating these internal signals, often outperforming traditional prompting methods. However, the effectiveness of this approach varies significantly across different models and attributes, highlighting the complexity of controlling implicit personalization.
Internal activation signals in LLMs can be manipulated to suppress implicit demographic influences more effectively than traditional prompting methods.
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.