Search papers, labs, and topics across Lattice.
This study reveals that the effective learning rate (ELR) is a critical factor governing loss dynamics in language model pretraining, as it allows for the collapse of loss trajectories when matched across different runs, regardless of variations in learning rates and parameter norms. Systematic ablations indicate that normalization design and the timescale of learning rate and norm variations significantly influence the precision of this collapse. The research also introduces a fitted functional scaling law (FSL) that leverages ELR to unify norm-control methods, providing insights into delayed acceleration effects in training dynamics.
Matching effective learning rates across different training setups leads to remarkably consistent loss trajectories, challenging conventional wisdom about learning rate variability.
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few x 10^-3, below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variation as key determinants of collapse precision. Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce. Replacing LR with ELR enables a fitted functional scaling law (FSL) to transfer across norm-control methods. The resulting ELR-based FSL also explains delayed acceleration, a recurring effect of norm control. Together, these results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics.