Search papers, labs, and topics across Lattice.
This study introduces Group Alignment-induced Sycophancy (GAS), a framework that evaluates the dual impact of group alignment on language models by measuring both the alignment with demographic group opinions and the resulting sycophantic behavior. The authors systematically assess three alignment methods across four models and thirteen demographic groups, revealing that the gains in opinion alignment and shifts in sycophancy are not uniform and vary significantly between groups. This nuanced understanding suggests that alignment evaluations should consider a multi-dimensional profile rather than a singular score, highlighting the complexity of adapting LLMs for diverse populations.
Group alignment can lead to unexpected sycophantic behavior, with some demographic groups experiencing greater alignment gains than others, challenging the notion of a one-size-fits-all approach in model adaptation.
Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and objective information. However, existing group alignment methods and evaluations focus only on how closely the model matches the group's opinions, overlooking the induced change in sycophantic behaviour. To bridge this gap, we introduce \textbf{G}roup \textbf{A}lignment-induced \textbf{S}ycophancy (GAS) and systematically evaluate alignment across 3 methods, 4 models and 13 demographic groups, on both the intended gain in opinion alignment and the unintended shift in sycophancy. We find that gain and shift are non-uniform across groups: under an identical budget, some groups receive larger gains in opinion alignment than others, and the induced sycophancy shift forms a group-specific profile rather than a single-dimensional change. These results suggest that group alignment should be reported as a two-sided, multi-dimensional profile rather than a single fit score that accounts for per-group differences when adapting LLMs to diverse populations.