Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
2
0
1
Excessive weight decay can destabilize training by driving scale-invariant weight norms past a critical boundary, leading to sudden loss spikes.
Adam can achieve linear convergence on highly degenerate polynomials without careful tuning, thanks to a built-in mechanism that exponentially amplifies the effective learning rate.