Search papers, labs, and topics across Lattice.
This paper introduces a neighboring early-stopping rule for adaptive regularization in kernel ridge regression with random features (KRR-RF), addressing the challenge of selecting the optimal regularization parameter without prior knowledge of smoothness and capacity parameters. By utilizing a grid that focuses on adjacent estimators, the method significantly reduces the number of necessary discrepancy comparisons compared to traditional all-pairs approaches. The authors establish that their method achieves oracle polynomial learning rates under standard conditions, demonstrating improved prediction performance and computational efficiency in simulations and real-data applications.
Achieving oracle-rate guarantees without prior knowledge of model parameters could revolutionize how we approach regularization in kernel methods.
Random feature methods provide a scalable approximation to kernel ridge regression (KRR), but the regularization parameter that yields the oracle learning rate depends on unknown smoothness and capacity parameters. In this work, we propose a neighboring early-stopping rule for adaptive regularization in KRR with random features (KRR-RF). The method uses a grid that is uniform in inverse regularization and compares only adjacent estimators, reducing the number of discrepancy comparisons relative to standard all-pairs Lepskii-type procedures. Both the neighboring discrepancy and its empirical complexity term can be computed directly in the random feature space, without constructing the exact kernel Gram matrix. We establish a high-probability comparison bound for neighboring KRR-RF estimators and show that, under standard source and capacity conditions together with suitable grid and random feature budget conditions, the selected estimator attains the oracle polynomial learning rate up to logarithmic factors. The result allows the regularization parameter to be selected without prior knowledge of the source and capacity exponents and covers both well-specified and partially misspecified regimes. Our analysis is based on an empirical random feature effective dimension that connects the observable stopping threshold with the population complexity of the random feature model. Simulation and real-data experiments illustrate the prediction performance and computational behavior of the proposed method in comparison with standard tuning procedures.