Search papers, labs, and topics across Lattice.
By formalizing stochastic gradient descent as a Markovian stochastic process, the authors derive an information-theoretic speed limit that fundamentally bounds the rate at which model parameters acquire Fisher information about latent data-generating factors. This formulation decomposes the information flow into deterministic drift forces and stochastic gradient fluctuations, isolating the exact dynamical contributions to representation emergence. Validated on basis-function linear regression, the resulting bound accurately predicts both the characteristic timescales and the sequential order in which latent variables are encoded during training.
Neural network feature learning obeys a strict Fisher-information speed limit, governing the exact sequence and timescales with which latent representations can emerge from stochastic gradient dynamics.
Neural networks acquire internal representations through learning. In this work, we formulate stochastic gradient descent (SGD) as a Markovian stochastic process and derive a Fisher-information flow speed limit that bounds the rate at which trainable parameters can acquire information about latent variables in the data-generating process. The resulting inequality decomposes the information flow into drift and noise contributions, thereby quantifying the roles of deterministic learning forces and SGD-induced fluctuations from an information-theoretic perspective. We verify the bound in analytically tractable basis-function linear regression, where the information budget predicted by the bound reproduces the ordering and characteristic time scales with which different latent variables are encoded in the learned parameters. These results establish Fisher-information speed limits as a quantitative framework for diagnosing when and how different aspects of the data-generating mechanism are acquired during stochastic learning.