Search papers, labs, and topics across Lattice.
This study investigates the propagation of various input perturbations through decoder-only language models, focusing on six naturalistic and synthetic types of disruptions. By analyzing output behavior, hidden-state geometry, and attention-head function across multiple checkpoints of GPT-2 and Qwen2.5, the authors reveal that perturbations create distinct metric profiles that are inadequately captured by traditional output measures. The findings indicate that robustness assessments based on a singular metric can be misleading, highlighting the necessity for a multi-level evaluation approach to understand how perturbations impact language model computations.
Robustness claims in language models can be misleading when based solely on output behavior, as perturbations reveal complex, multi-level effects that traditional metrics overlook.
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints by analyzing layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores are especially associated with activation-patching recovery under token substitution and shuffling. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.