Search papers, labs, and topics across Lattice.
This study investigates how scaling model-generated distillation data enhances the recovery of latent teacher traits in student models, revealing that larger datasets can amplify subtle signals from teachers even when the data is off-task. By employing a controlled subliminal learning setup, the researchers demonstrate that students trained on varying amounts of independent off-task data show a clearer expression of the targeted traits, particularly when the initial student model is less aligned with the desired behavior. The findings indicate that not only does scaling improve trait detection, but it can also shift student behavior towards intended traits when the initial model favors alternatives, suggesting a nuanced relationship between data scale and trait expression.
Larger datasets can reveal hidden teacher traits in student models, even from off-task data, potentially reshaping how we approach model distillation.
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.