Search papers, labs, and topics across Lattice.
This paper investigates the gradient flow dynamics of diagonal linear networks under infinitesimal initialization, extending previous results to encompass both deep and two-layer architectures. The authors introduce a novel algorithm that characterizes the training trajectories of these networks and prove its convergence to a modified \( \mathcal{l}_1 \) norm minimization solution. Key findings reveal that the implicit bias of these networks aligns with this modified \( \mathcal{l}_1 \) norm, with the Structural Invariant Manifold identified as a crucial geometric factor influencing the learning dynamics.
The implicit bias of diagonal linear networks under infinitesimal initialization reveals a surprising alignment with a modified \( \mathcal{l}_1 \) norm, reshaping our understanding of their training dynamics.
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.1). Specifically, we demonstrate that the training trajectories of these models can be equivalently characterized by the proposed Algorithm 1. We further prove that this algorithm converges to the solution of a modified $ \mathcal{l}_1 $ norm minimization problem. As a result, we establish that the implicit bias of both network architectures corresponds to a modified $ \mathcal{l}_1 $ norm in the regime of infinitesimal initialization. Additionally, we provide insights into the underlying mechanisms governing these dynamics by identifying the Structural Invariant Manifold (SIM) (Zhao et al., 2026) as the key geometric structure that shapes the learning process.