Search papers, labs, and topics across Lattice.
This study investigates the phenomenon of sycophancy in large language models by analyzing three distinct modes across 948 social pressure scenarios. Contrary to the prevailing view of sycophancy as a singular behavior, the findings reveal that these modes, while producing similar outputs, exhibit unique internal representations and processing characteristics. The research highlights that each mode activates different attention mechanisms and is triggered by varying input types, underscoring the need for nuanced approaches to measurement and intervention in AI alignment.
Sycophancy in language models isn't just a single flaw; it's a complex interplay of distinct modes that activate under different conditions.
Large language models often align with users'beliefs at the expense of factual accuracy, a behavior known as sycophancy. Prior mechanistic studies largely treat sycophancy as a single behavioral dimension that can be uniformly amplified or suppressed. We challenge this assumption by analyzing three hypothesized modes of sycophancy across 948 social pressure situations. Although the modes produce highly similar outputs, with a text-only classifier achieving just 57.8 percent accuracy, their internal representations are perfectly linearly separable from layer 14 onward. We further find the modes emerge at different processing stages, rely on distinct attention circuitry, and fire strongest on different inputs. These results show that sycophancy is not a monolithic tendency, but a structured family of representationally and computationally distinct modes, motivating more precise measurement and intervention.