Search papers, labs, and topics across Lattice.
3
0
3
Local SGD's efficiency can be rigorously explained through bounded second-order heterogeneity, revealing nearly tight convergence rates for general convex objectives.
Forget tedious hyperparameter tuning: adaptive optimizers like DP-SignSGD and DP-Adam maintain performance across privacy levels, unlike DP-SGD whose learning rate plummets with increased privacy.
AdamW's decoupled weight decay prevents Neural Collapse, challenging the assumption that this phenomenon is universal across optimization methods.