Search papers, labs, and topics across Lattice.
This paper analyzes a variant of stochastic gradient descent with initial regularization (SGDIR) and derives dimension-free upper bounds on its expected excess risk for squared loss, revealing significant improvements in risk bounds under specific conditions. In the noiseless scenario, the authors present new bounds that scale as \(m^{-2}\log^{2}m\) and \(m^{-3+\epsilon}\) for different parameter settings, while also establishing a matching lower bound in certain regimes. In the noisy case, SGDIR is shown to have expected excess risk comparable to ridge regression, providing theoretical support for its efficacy in practical applications.
SGDIR achieves superior risk bounds compared to traditional methods, challenging the effectiveness of ridge regression under noise.
We analyze a variant of stochastic gradient descent with initial regularization (SGDIR) and derive dimension-free upper bounds on its expected excess risk for the squared loss. In the noiseless case, we obtain new bounds for both averaged and non-averaged SGDIR under moment, source, and capacity assumptions. For a particular value of the source parameter, these bounds are of order $m^{-2}\log^{2}m$, where the number of training samples is of order $m$. For another value of the source parameter, we obtain, for any $蔚>0$, bounds of order $m^{-3+蔚}$, provided that the capacity parameter exceeds $蔚^{-1}$. We also establish a lower bound that matches our upper bounds in certain regimes up to a polylogarithmic factor. In the noisy case, we provide an instance-based comparison between SGDIR and ridge regression. Under general assumptions and a mild lower bound on the regularization parameter, we show that the expected excess risk of SGDIR is no larger than that of ridge regression, up to a polylogarithmic factor. Numerical experiments on synthetic and real data are consistent with our theoretical findings.