Search papers, labs, and topics across Lattice.
This paper investigates the relationship between gradient descent (GD) and gradient flow in the context of ReLU training, revealing that differentiation of GD does not commute with the continuous-time limit. The authors establish that while GD states converge over a fixed horizon, the discrete derivatives approach a regional propagator that incorporates activation-event transfers, leading to significant discrepancies in curvature. Notably, they demonstrate that standard residual-ReLU models can exhibit arbitrarily large sensitivity ratios, challenging assumptions about stability in initialization sets.
Discrete differentiation of gradient descent reveals unexpected curvature behaviors that challenge conventional stability assumptions in ReLU training.
Gradient descent (GD) is explicit Euler for gradient flow, but a state-accurate continuous-time surrogate need not remain accurate after differentiation. At every fixed nonresonant step size, ordinary automatic differentiation exactly differentiates the executed hard-ReLU GD program. We prove that, over a fixed finite horizon, the GD states converge and these exact discrete derivatives approach an event-free regional propagator, whereas the derivative of the limiting flow also contains speed-normalized activation-event transfers. A prepoint Stieltjes representation separates the absolutely continuous regional Hessian from atomic interface curvature; one nonzero gradient jump produces an exactly rank-one endpoint discrepancy, and global convexity prevents complete multi-event cancellation whenever an event is strict. Nevertheless, a standard family of globally 1-strongly convex residual-ReLU squared-loss risks realizes arbitrarily large reciprocal sensitivity ratios on open initialization sets, with a uniform transversality margin. The same discrete-versus-flow decomposition extends to parameters and reverse-mode adjoints; resolved smoothing in the scalar or autonomous-normal regime and consistent event localization recover the flow sensitivity. The results concern deterministic full-batch, finite-horizon dynamics with a stable finite itinerary of separated same-direction transverse events; they are consistency theorems, not prevalence claims for large-scale training.