Search papers, labs, and topics across Lattice.
To explain how stochastic optimizers collapse overparameterized neural networks into simpler subnetworks, this work models stochastic gradient flow as a percolation process governed by architectural symmetries. The authors find that subnetworks coalesce in discrete, simultaneous clusters rather than continuously, manifesting as sharp variance spikes in a macroscopic order parameter characteristic of physical phase transitions. This condensation mechanism and its discrete scale-invariant cascades are shown to generalize from SGD to adaptive optimizers like Adam and AdamW under heavy-tailed gradient noise.
Neural networks prune themselves during optimization not through gradual simplification, but via abrupt percolation phase transitions where architectural symmetries force subnetworks to merge in discrete, cascading blocks.
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.