Search papers, labs, and topics across Lattice.
This paper introduces Cross-Encoding Steering Evaluation to analyze how activation steering influences answer encodings and extraction indices. The authors demonstrate that contrastive activation addition (CAA) leads to significant score changes favoring extraction indices over semantic labels, particularly at deeper layers of the model. Their findings reveal that steering gains are not universally indicative of intervention control, highlighting the importance of context in evaluating model behavior across different answer encodings.
Extraction indices, not semantic labels, dominate the influence of activation steering, reshaping our understanding of model behavior across varying contexts.
Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during construction. We introduce Cross-Encoding Steering Evaluation, which freezes an intervention while re-encoding answers to the same held-out items. On NormBank, after A/B/C identifiers are reassigned, contrastive activation addition (CAA) induces larger target-versus-source score changes for the extraction indices than for the semantic labels under the new mapping. We call this extraction-index following. Varying identifier vocabulary (A/B/C, X/Y/Z, or 1/2/3) and row order shows that the effect tracks extraction index rather than row position. After matching direction norms across layers, extraction-index following emerges mainly at later depths. A low-rank output-sensitive component containing 15.4% of the direction's squared norm retains 96.3% of this effect. An Inference-Time Intervention (ITI)-style method also favors extraction-index over semantic-label following on NormBank in three models. In aggregate, MNLI favors extraction-index following, whereas Social Chemistry 101 (SC101) favors semantic-label following. Multiple-choice and open-ended evaluations can yield different behavioral conclusions. Thus, a steering gain under one answer encoding does not by itself identify what the intervention controls.