Search papers, labs, and topics across Lattice.
This paper introduces SteerCheck, an attribution audit tool designed to assess the specificity of activation steering effects in AI models. By analyzing 960 interventions on the Qwen3-14B model, the authors reveal that while isotropic directions are limited, sign-randomized directions often maintain significant alignment with target concepts, evidenced by a strong correlation with signed cosine metrics. The findings highlight the importance of understanding alignment leakage and its implications for the reliability of conditional randomization tests in model evaluations.
Activation steering can mislead evaluations, with over 25% of interventions showing unexpected alignment leakage that complicates model audits.
Activation steering can change behaviour without establishing that the effect is specific to the intended concept. We introduce SteerCheck, a preregistered attribution audit that matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims. Exact replay of 960 Qwen3-14B interventions reveals complementary limits of common controls: isotropic directions occupy a narrow near-orthogonal region, whereas sign-randomized same-construction directions often retain substantial target alignment. Effect is strongly associated with signed cosine within the sign-randomized family ($蟻=.94$); $25.3\%$ of its draws exceed cosine $.5$, and every draw exceeding the observed mean effect has cosine above $.80$. This alignment leakage does not by itself invalidate a conditional randomization test; it limits what the comparator can distinguish and motivates reporting exchangeability assumptions, a construction diagnostic $A$, and the empirical cosine distribution. The primary Qwen complete gate remains negative because the protected tail fails all families. On independent data, continuous margin transfers only in Qwen and accuracy transfers in no selected cell. Prospectively registered language controls pass the complete gate in Qwen and DeepSeek, while a passing DeepSeek detox comparator rules out categorical separation; all nominal passes are sensitive to $螕=1.10$. Frozen three-rater open-generation evaluation supports factual correction in DeepSeek but not Qwen; the automatic judge fails calibration (macro-F1 $.562$), so null-wide semantic results remain descriptive. SteerCheck makes these conditional and mixed conclusions auditable.