Search papers, labs, and topics across Lattice.
2
0
1
Vanilla gradient descent cannot beat the silver stepsize schedule, proving that predetermined learning rates hit a hard theoretical limit of $\mathcal{O}(n^{-\log_2(1+\sqrt{2})})$ without momentum.
Achieving a new anytime lower bound of \(惟(n^{-1.2408})\) reveals critical insights into the acceleration limits of gradient descent.