Search papers, labs, and topics across Lattice.
This study investigates the dual behaviors of sycophancy in large language models, distinguishing between Unsupported-Yielding and Rational-Updating in response to user feedback. Using a two-turn evaluation framework, the authors demonstrate that interventions aimed at reducing sycophancy can inadvertently impair the model's ability to rationally update its responses based on valid user input. Mechanistic analysis reveals significant overlap in the neural substrates driving both behaviors, suggesting that effective anti-sycophancy strategies must balance the suppression of Unsupported-Yielding with the preservation of Rational-Updating capabilities.
Reducing sycophancy in language models can inadvertently hinder their ability to rationally update, revealing a critical trade-off in model behavior.
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains useful evidence, prompting the model to update its answer in a rational way. We distinguish them as Unsupported-Yielding and Rational-Updating. Prior work focuses primarily on suppressing Unsupported-Yielding, while overlooking its effect on Rational-Updating. We address this gap with a two-turn evaluation framework that measures the two behaviors separately. Across representative training-time and inference-time interventions, we find that anti-sycophancy methods often encounter a trade-off in which reducing Unsupported-Yielding can sacrifice Rational-Updating, and vice versa, even when the two objectives are optimized jointly. Mechanistic analysis suggests that the two behaviors share an internal substrate: the MLP neurons and attention heads driving them overlap substantially, and their associated steering directions are positively aligned. We further conduct a preliminary orthogonalized steering exploration, which yields modest, backbone-dependent selectivity gains. Overall, our results suggest that anti-sycophancy should be treated not as a simple suppression problem, but as a selectivity problem, where effective interventions should preserve Rational-Updating while reducing Unsupported-Yielding.