Search papers, labs, and topics across Lattice.
This paper introduces the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), which effectively samples from non-convex target distributions with non-smooth potentials and superlinear gradient growth by directly utilizing subgradients. The method employs taming techniques to ensure stability and derives non-asymptotic convergence bounds in Wasserstein-2 distance, significantly enhancing the performance of subgradient-based Langevin algorithms. Experimental results demonstrate that SG-TULA competes favorably with traditional optimization methods like AdamW and Muon, providing explicit guarantees that are currently lacking in the field.
SG-TULA achieves competitive performance in sampling from complex non-convex distributions while offering explicit convergence guarantees that traditional methods fail to provide.
We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.