Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
arXiv:1611.01838
Abstract
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenvalues. We leverage upon this observation to construct a local-entropy-based objective function that favors well-generalizable solutions lying in large flat regions of the energy landscape, while avoiding poorly-generalizable solutions located in the sharp valleys. Conceptually, our algorithm resembles two nested loops of SGD where we use Langevin dynamics in the inner loop to compute the gradient of the local entropy before each update of the weights. We show that the new objective has a smoother energy landscape and show improved generalization over SGD using uniform stability, under certain assumptions. Our experiments on convolutional and recurrent networks demonstrate that Entropy-SGD compares favorably to state-of-the-art techniques in terms of generalization error and training time.
ICLR '17
References in corpus (17)
- Striving for Simplicity: The All Convolutional Net
- Recurrent Neural Network Regularization
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- Understanding deep learning requires rethinking generalization
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Qualitatively characterizing neural network optimization problems
- No bad local minima: Data independent training error guarantees for multilayer neural networks
- Efficient approaches for escaping higher order saddle points in non-convex optimization
- Bridging the Gap between Stochastic Gradient MCMC and Stochastic Optimization
- Local entropy as a measure for sampling solutions in Constraint Satisfaction Problems
- A Variational Analysis of Stochastic Gradient Algorithms
- Training Recurrent Neural Networks by Diffusion
- Learning may need only a few bits of synaptic precision
- On the energy landscape of deep networks
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
Cited by in corpus (72)
- Don't Decay the Learning Rate, Increase the Batch Size
- GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration
- Knowledge Distillation by On-the-Fly Native Ensemble
- A Closer Look at Memorization in Deep Networks
- Exploring Generalization in Deep Learning
- Don't Use Large Mini-Batches, Use Local SGD
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks
- Optimization for deep learning: theory and algorithms
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- To understand deep learning we need to understand kernel learning
- Understanding and Enhancing the Transferability of Adversarial Examples
- Stronger generalization bounds for deep nets via a compression approach
- Implicit Regularization in Deep Learning
- Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
- Newton-Type Methods for Non-Convex Optimization Under Inexact Hessian Information
- Self-Augmentation: Generalizing Deep Networks to Unseen Classes for Few-Shot Learning
- Towards Binary-Valued Gates for Robust LSTM Training
- Understanding Batch Normalization
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- The loss landscape of overparameterized neural networks
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- Implicit Bias of Gradient Descent on Linear Convolutional Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Classification regions of deep neural networks
- Laplacian Smoothing Gradient Descent
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
- Convergent Block Coordinate Descent for Training Tikhonov Regularized Deep Neural Networks
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- A Scale Invariant Flatness Measure for Deep Network Minima
- Parle: parallelizing stochastic gradient descent
- Stochastic Training of Residual Networks: a Differential Equation Viewpoint
- Emergent properties of the local geometry of neural loss landscapes
- Stagewise Training Accelerates Convergence of Testing Error Over SGD
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- Information Bottleneck and its Applications in Deep Learning
- A Stochastic Composite Gradient Method with Incremental Variance Reduction
- Simulated Tempering Langevin Monte Carlo II: An Improved Proof using Soft Markov Chain Decomposition
- LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- Selection dynamics for deep neural networks
- Uniform-in-Time Weak Error Analysis for Stochastic Gradient Descent Algorithms via Diffusion Approximation
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- The Impact of Local Geometry and Batch Size on Stochastic Gradient Descent for Nonconvex Problems
- The asymptotic spectrum of the Hessian of DNN throughout training
- Consistent Sparse Deep Learning: Theory and Computation
- How neural networks find generalizable solutions: Self-tuned annealing in deep learning
- Robust and On-the-fly Dataset Denoising for Image Classification
- Post-synaptic potential regularization has potential
- Improved Techniques For Weakly-Supervised Object Localization
- Reinforced stochastic gradient descent for deep neural network learning
- Gradient Energy Matching for Distributed Asynchronous Gradient Descent
- Fast, Better Training Trick -- Random Gradient
- Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks
- Deep Curvature Suite
- Heavy-ball Algorithms Always Escape Saddle Points
- SaaS: Speed as a Supervisor for Semi-supervised Learning
- The sharp, the flat and the shallow: Can weakly interacting agents learn to escape bad minima?
- Information Theoretic Interpretation of Deep learning
- Optimization of neural networks via finite-value quantum fluctuations
- Nonlinear Collaborative Scheme for Deep Neural Networks
- TensOrMachine: Probabilistic Boolean Tensor Decomposition
- Non-Convex Optimization with Spectral Radius Regularization
- Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent
- Stochastic Backward Euler: An Implicit Gradient Descent Algorithm for -means Clustering
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
- TATi-Thermodynamic Analytics ToolkIt: TensorFlow-based software for posterior sampling in machine learning applications
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- Sequenced-Replacement Sampling for Deep Learning