Search papers, labs, and topics across Lattice.
This study investigates the phenomenon of weird generalization (WG) in language models, revealing that its occurrence is significantly influenced by the composition and language of fine-tuning datasets rather than their size. Through experiments with three open-weight models across four datasets, the authors found that WG is more pronounced when familiar data from pretraining is used and is highly sensitive to the specific evaluation questions posed. These findings suggest that WG arises from fragile interactions between training and evaluation data, indicating that it poses an adversarial threat that necessitates careful data engineering rather than being a routine risk of fine-tuning.
Weird generalization is not just a quirk of model training; it鈥檚 a fragile phenomenon that could be weaponized if not carefully managed.
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.