Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
arXiv:1512.07666
Abstract
Effective training of deep neural networks suffers from two main issues. The first is that the parameter spaces of these models exhibit pathological curvature. Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD). These methods improve convergence by adapting to the local geometry of parameter space. A second issue is overfitting, which is typically addressed by early stopping. However, recent work has demonstrated that Bayesian model averaging mitigates this problem. The posterior can be sampled by using Stochastic Gradient Langevin Dynamics (SGLD). However, the rapidly changing curvature renders default SGLD methods inefficient. Here, we propose combining adaptive preconditioners with SGLD. In support of this idea, we give theoretical properties on asymptotic convergence and predictive risk. We also provide empirical results for Logistic Regression, Feedforward Neural Nets, and Convolutional Neural Nets, demonstrating that our preconditioned SGLD method gives state-of-the-art performance on these models.
AAAI 2016
References in corpus (9)
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- Weight Uncertainty in Neural Networks
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- On the Convergence of Stochastic Gradient MCMC Algorithms with High-Order Integrators
- Bayesian Posterior Sampling via Stochastic Gradient Fisher Scoring
- Consistency and fluctuations for stochastic gradient Langevin dynamics
Cited by in corpus (76)
- Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
- Preconditioned Stochastic Gradient Descent
- Bridging the Gap between Stochastic Gradient MCMC and Stochastic Optimization
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- Stein Variational Gradient Descent With Matrix-Valued Kernels
- Deep Latent Dirichlet Allocation with Topic-Layer-Adaptive Stochastic Gradient Riemannian MCMC
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- The True Cost of Stochastic Gradient Langevin Dynamics
- Stochastic Quasi-Newton Langevin Monte Carlo
- Structured Variational Learning of Bayesian Neural Networks with Horseshoe Priors
- Stochastic Gradient MCMC with Repulsive Forces
- Riemannian Stein Variational Gradient Descent for Bayesian Inference
- Meta-Learning for Stochastic Gradient MCMC
- Finding Mixed Nash Equilibria of Generative Adversarial Networks
- Bayesian Inference for Large Scale Image Classification
- Weighted -contractivity of Langevin dynamics with singular potentials
- DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference
- Inexact Newton Methods for Stochastic Nonconvex Optimization with Applications to Neural Network Training
- Robust Reinforcement Learning via Adversarial training with Langevin Dynamics
- Contextual Dropout: An Efficient Sample-Dependent Dropout Module
- Icebreaker: Element-wise Active Information Acquisition with Bayesian Deep Latent Gaussian Model
- Bayesian Graph Convolutional Neural Networks Using Non-Parametric Graph Learning
- EVA: Generating Longitudinal Electronic Health Records Using Conditional Variational Autoencoders
- On the Ergodicity, Bias and Asymptotic Normality of Randomized Midpoint Sampling Method
- A Bayesian Perspective on the Deep Image Prior
- Bayesian Sparse learning with preconditioned stochastic gradient MCMC and its applications
- Sampling-based Bayesian Inference with gradient uncertainty
- A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal Distributions
- Distributed Bayesian Learning with Stochastic Natural-gradient Expectation Propagation and the Posterior Server
- A Convergence Analysis for A Class of Practical Variance-Reduction Stochastic Gradient MCMC
- Scalable Natural Gradient Langevin Dynamics in Practice
- Generalized Bayesian Likelihood-Free Inference
- On the Generalised Langevin Equation for Simulated Annealing
- Non-Parametric Graph Learning for Bayesian Graph Neural Networks
- Multi-variance replica exchange stochastic gradient MCMC for inverse and forward Bayesian physics-informed neural network
- AMAGOLD: Amortized Metropolis Adjustment for Efficient Stochastic Gradient MCMC
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Differential Bayesian Neural Nets
- Data clustering based on Langevin annealing with a self-consistent potential
- An adaptive Hessian approximated stochastic gradient MCMC method
- A benchmark study on reliable molecular supervised learning via Bayesian learning
- Stochastic Particle-Optimization Sampling and the Non-Asymptotic Convergence Theory
- DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving
- When in Doubt: Neural Non-Parametric Uncertainty Quantification for Epidemic Forecasting
- Differentially Private Bayesian Neural Networks on Accuracy, Privacy and Reliability
- Non-Convex Optimization via Non-Reversible Stochastic Gradient Langevin Dynamics
- Stochastic Gradient Langevin Dynamics Algorithms with Adaptive Drifts
- Accelerating Convergence of Replica Exchange Stochastic Gradient MCMC via Variance Reduction
- Mini-batch Metropolis-Hastings MCMC with Reversible SGLD Proposal
- Characterizing Membership Privacy in Stochastic Gradient Langevin Dynamics
- Adaptively Preconditioned Stochastic Gradient Langevin Dynamics
- Stochastic natural gradient descent draws posterior samples in function space
- Probabilistic Approximate Logic and its Implementation in the Logical Imagination Engine
- Sampling with Mirrored Stein Operators
- Disentangling the Roles of Curation, Data-Augmentation and the Prior in the Cold Posterior Effect
- High-Order Stochastic Gradient Thermostats for Bayesian Learning of Deep Models
- On Connecting Stochastic Gradient MCMC and Differential Privacy
- Laplacian Smoothing Stochastic Gradient Markov Chain Monte Carlo
- Deep Poisson gamma dynamical systems
- Deep Autoencoding Topic Model with Scalable Hybrid Bayesian Inference
- Stochastic Gradient MCMC with Stale Gradients
- Variationally Inferred Sampling Through a Refined Bound for Probabilistic Programs
- On Transformations in Stochastic Gradient MCMC
- A Langevinized Ensemble Kalman Filter for Large-Scale Static and Dynamic Learning
- Isotropic SGD: a Practical Approach to Bayesian Posterior Sampling
- A fast asynchronous MCMC sampler for sparse Bayesian inference
- Secure and Differentially Private Bayesian Learning on Distributed Data
- An Adaptive Empirical Bayesian Method for Sparse Deep Learning
- MCMC-Interactive Variational Inference
- Efficient and Transferable Adversarial Examples from Bayesian Neural Networks
- Contributions to Large Scale Bayesian Inference and Adversarial Machine Learning
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- Approximate Inference via Clustering
- A Divergence Bound for Hybrids of MCMC and Variational Inference and an Application to Langevin Dynamics and SGVI
- Thompson Sampling via Local Uncertainty
- Structured Stochastic Gradient MCMC