An empirical analysis of the optimization of deep network loss surfaces
arXiv:1612.04010
Abstract
The success of deep neural networks hinges on our ability to accurately and efficiently optimize high-dimensional, non-convex functions. In this paper, we empirically investigate the loss functions of state-of-the-art networks, and how commonly-used stochastic gradient descent variants optimize these loss functions. To do this, we visualize the loss function by projecting them down to low-dimensional spaces chosen based on the convergence points of different optimization algorithms. Our observations suggest that optimization algorithms encounter and choose different descent directions at many saddle points to find different final weights. Based on consistency we observe across re-runs of the same stochastic optimization algorithm, we hypothesize that each optimization algorithm makes characteristic choices at these saddle points.
References in corpus (5)
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Qualitatively characterizing neural network optimization problems
- No bad local minima: Data independent training error guarantees for multilayer neural networks
- Deep Learning without Poor Local Minima
- Mean field theory of spin glasses: statics and dynamics
Cited by in corpus (15)
- Visualizing the Loss Landscape of Neural Nets
- A Closer Look at Memorization in Deep Networks
- How Does Batch Normalization Help Optimization?
- U(1) symmetric recurrent neural networks for quantum state reconstruction
- Data efficiency and extrapolation trends in neural network interatomic potentials
- Interpreting Adversarial Robustness: A View from Decision Surface in Input Space
- Exploring loss function topology with cyclical learning rates
- Multi-class Classification without Multi-class Labels
- Quantitatively Evaluating GANs With Divergences Proposed for Training
- Demystifying Batch Normalization in ReLU Networks: Equivalent Convex Optimization Models and Implicit Regularization
- On the Distributional Properties of Adaptive Gradients
- Model-Agnostic Meta-Learning using Runge-Kutta Methods
- SPI-Optimizer: an integral-Separated PI Controller for Stochastic Optimization
- Traversing the noise of dynamic mini-batch sub-sampled loss functions: A visual guide
- The layer-wise L1 Loss Landscape of Neural Nets is more complex around local minima