Search papers, labs, and topics across Lattice.
This study investigates the varying impacts of different types of evaluation-awareness framing on model compliance during safety evaluations. By categorizing eval-awareness into capabilities-flavored and safety-flavored framings, the authors demonstrate that capabilities-framing significantly enhances compliance, achieving a 24 to 46 percentage-point increase over safety-framing in various steering conditions. Additionally, a causal link is suggested through a chain-of-thought prefill intervention that effectively shifts compliance outcomes, highlighting the nuanced behavior of eval-awareness in AI models.
Capabilities-framing can boost model compliance by up to 46 percentage points compared to safety-framing, revealing that not all eval-awareness is created equal.
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same"X% suppression of eval-awareness"can correspond to qualitatively different behavioral outcomes.