Search papers, labs, and topics across Lattice.
Across more than 2,400 pre-training runs spanning 271M to 8.2B parameters, this study systematically rehabilitates layer dropout (stochastic depth) for modern LLM pre-training by optimizing its layer distribution, temporal schedules, and optimizer hyperparameters. When properly scheduled, layer dropout yields strictly lower validation loss at iso-FLOPs or saves up to 25% of training FLOPs at iso-steps compared to standard dense baselines. Crucially, the induced structural robustness transfers directly to post-training inference, enabling up to 1.5x speedups via zero-shot layer skipping, early exiting, and self-speculative decoding with negligible quality degradation.
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, we observe that at the same training FLOPs layer dropout leads to lower loss. For a given number of training steps, LLMs can achieve lower or similar validation loss while saving upto 25% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.5x inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.