Search papers, labs, and topics across Lattice.
2
0
2
The implicit bias of diagonal linear networks under infinitesimal initialization reveals a surprising alignment with a modified \( \mathcal{l}_1 \) norm, reshaping our understanding of their training dynamics.
Adam can achieve linear convergence on highly degenerate polynomials without careful tuning, thanks to a built-in mechanism that exponentially amplifies the effective learning rate.