Search papers, labs, and topics across Lattice.
This study investigates the phenomenon of grokking in overparameterized neural networks, focusing on how weight-decay (WD) pulses influence generalization during training. By systematically applying short WD pulses, the authors reveal a dose-ordered timing effect where stronger WD increases lead to earlier generalization, while decreases delay it, all occurring before observable generalization manifests. This research highlights the concept of canalization in function selection, providing insights into the dynamics of solution selection in neural networks.
Weight-decay pulses can dictate the timing of generalization in neural networks, revealing a surprising canalization effect that precedes visible performance improvements.
For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-decay (WD) pulses across this plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. This ordering emerges before visible generalization in all three tasks. Meanwhile, test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. We call this combination of increasingly constrained solution selection and persistent dose-ordered timing sensitivity the canalization of function selection.