Bridging the Gap between Stochastic Gradient MCMC and Stochastic Optimization
arXiv:1512.07962
Abstract
Stochastic gradient Markov chain Monte Carlo (SG-MCMC) methods are Bayesian analogs to popular stochastic optimization methods; however, this connection is not well studied. We explore this relationship by applying simulated annealing to an SGMCMC algorithm. Furthermore, we extend recent SG-MCMC methods with two key components: i) adaptive preconditioners (as in ADAgrad or RMSprop), and ii) adaptive element-wise momentum weights. The zero-temperature limit gives a novel stochastic optimization method with adaptive element-wise momentum weights, while conventional optimization methods only have a shared, static momentum weight. Under certain assumptions, our theoretical analysis suggests the proposed simulated annealing approach converges close to the global optima. Experiments on several deep neural network models show state-of-the-art results compared to related stochastic optimization algorithms.
Merry Christmas from the Santa (algorithm). AISTATS 2016
References in corpus (8)
- ADADELTA: An Adaptive Learning Rate Method
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Modeling Temporal Dependencies in High-Dimensional Sequences: Application to Polyphonic Music Generation and Transcription
- Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
- On the Convergence of Stochastic Gradient MCMC Algorithms with High-Order Integrators
- Consistency and fluctuations for stochastic gradient Langevin dynamics
- High-Order Stochastic Gradient Thermostats for Bayesian Learning of Deep Models
Cited by in corpus (21)
- Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
- Stochastic Gradient Descent as Approximate Bayesian Inference
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- A Variational Analysis of Stochastic Gradient Algorithms
- Predictive Coarse-Graining
- Global Convergence of Stochastic Gradient Hamiltonian Monte Carlo for Non-Convex Stochastic Optimization: Non-Asymptotic Performance Bounds and Momentum-Based Acceleration
- A Bayesian Inference Framework for Procedural Material Parameter Estimation
- Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMC
- Stochastic Gradient MCMC with Repulsive Forces
- Decentralized Stochastic Gradient Langevin Dynamics and Hamiltonian Monte Carlo
- Meta-Learning for Stochastic Gradient MCMC
- Icebreaker: Element-wise Active Information Acquisition with Bayesian Deep Latent Gaussian Model
- Asymptotic Analysis via Stochastic Differential Equations of Gradient Descent Algorithms in Statistical and Computational Paradigms
- Sampling-based Bayesian Inference with gradient uncertainty
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Stochastic Gradient Monomial Gamma Sampler
- Decentralized Langevin Dynamics
- Isotropic SGD: a Practical Approach to Bayesian Posterior Sampling
- Dense Uncertainty Estimation via an Ensemble-based Conditional Latent Variable Model
- Dictionary Learning Strategies for Compressed Fiber Sensing Using a Probabilistic Sparse Model
- Approximate Inference via Clustering