Search papers, labs, and topics across Lattice.
This study critiques the conventional method of evaluating model performance by contrasting activations on evaluative versus non-evaluative prompts, revealing that the choice of prompt significantly influences the reported scores. By systematically varying prompts while holding task text constant, the authors demonstrate that the variance in reported scores is largely attributable to prompt selection rather than model characteristics. The findings indicate that a single-prompt design is inadequate for reliable model comparisons, necessitating a more robust approach to evaluation that accounts for the influence of prompt choice.
The choice of prompt can skew model evaluation scores, revealing that two conflicting studies can be reconciled by simply altering the prompt used.
A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose:"a prompt that announces an evaluation"is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single-prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.