Search papers, labs, and topics across Lattice.
4
0
4
14
Early stopping can transform gradient descent from a source of suboptimality into a strategy for achieving minimax-optimal classification performance in noisy settings.
AEW achieves optimal performance in expectation for model selection aggregation, revealing a critical phase transition that could redefine its application in statistical learning.
LLMs can waste compute on incorrect reasoning, but this work shows how to provably and practically stop them mid-generation when they're going astray.
Multipass SGD can suffer from suboptimal generalization if the preconditioner misaligns the geometry of the population risk curvature and gradient noise, leading to a worse effective dimension.