Nonasymptotic analysis of Stochastic Gradient Hamiltonian Monte Carlo under local conditions for nonconvex optimization
arXiv:2002.05465
Abstract
We provide a nonasymptotic analysis of the convergence of the stochastic gradient Hamiltonian Monte Carlo (SGHMC) to a target measure in Wasserstein-2 distance without assuming log-concavity. Our analysis quantifies key theoretical properties of the SGHMC as a sampler under local conditions which significantly improves the findings of previous results. In particular, we prove that the Wasserstein-2 distance between the target and the law of the SGHMC is uniformly controlled by the step-size of the algorithm, therefore demonstrate that the SGHMC can provide high-precision results uniformly in the number of iterations. The analysis also allows us to obtain nonasymptotic bounds for nonconvex optimization problems under local conditions and implies that the SGHMC, when viewed as a nonconvex optimizer, converges to a global minimum with the best known rates. We apply our results to obtain nonasymptotic bounds for scalable Bayesian inference and nonasymptotic generalization bounds.
Accepted to Journal of Machine Learning Research (JMLR), 2024, to appear
References in corpus (8)
- Bayesian Posterior Sampling via Stochastic Gradient Fisher Scoring
- Underdamped Langevin MCMC: A non-asymptotic analysis
- The promises and pitfalls of Stochastic Gradient Langevin Dynamics
- Is There an Analog of Nesterov Acceleration for MCMC?
- Global Convergence of Stochastic Gradient Hamiltonian Monte Carlo for Non-Convex Stochastic Optimization: Non-Asymptotic Performance Bounds and Momentum-Based Acceleration
- On stochastic gradient Langevin dynamics with dependent data streams: the fully non-convex case
- Higher Order Langevin Monte Carlo Algorithm
- Nonasymptotic estimates for Stochastic Gradient Langevin Dynamics under local conditions in nonconvex optimization
Cited by in corpus (5)
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- Decentralized Stochastic Gradient Langevin Dynamics and Hamiltonian Monte Carlo
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
- A proof of convergence for the gradient descent optimization method with random initializations in the training of neural networks with ReLU activation for piecewise linear target functions