Search papers, labs, and topics across Lattice.
This study investigates how language models prioritize between contextual information and stored knowledge when faced with conflicting inputs. By employing counterfactual experiments and estimating authority directions from agreement prompts, the authors reveal that steering models along these directions can significantly influence their source choice, achieving a 30-68% reproduction of authority-induced shifts. However, the findings also indicate that authority representations are largely task-dependent, with only 9% transferability across tasks compared to 57% for local task-specific directions.
Steering language models can shift their reliance on context versus memory, but this authority is surprisingly task-specific.
When contextual information conflicts with the knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. However, steering along a direction does not establish causality: whether the unedited model would naturally use that direction or whether the direction is reusable across tasks. We test these distinctions through counterfactual experiments in unambiguous settings. First, we estimate authority directions from agreement prompts, in which the context and parametric knowledge support the same answer. We then interchange naturally occurring coordinates along these directions between matched prompts that direct the model to prioritize either the supplied context or its parametric knowledge. Across Qwen, Llama, and OLMo models, this intervention reproduces 30-68% of the authority-induced shift in source choice, whereas matched controls reproduce almost none. To test cross-task reuse, we learn authority directions on two tasks separately and see that cross-task transferability closes only 9% of the authority gap while the local direction learned on the given task closes 57%. These results distinguish authority representation, causal use, and cross-task causal reuse, and suggest that authority computations may be task-dependent, rather than reusable across tasks.