The Loss Surfaces of Multilayer Networks
arXiv:1412.0233
Abstract
We study the connection between the highly non-convex loss function of a simple model of the fully-connected feed-forward neural network and the Hamiltonian of the spherical spin-glass model under the assumptions of: i) variable independence, ii) redundancy in network parametrization, and iii) uniformity. These assumptions enable us to explain the complexity of the fully decoupled neural network through the prism of the results from random matrix theory. We show that for large-size decoupled networks the lowest critical values of the random loss function form a layered structure and they are located in a well-defined band lower-bounded by the global minimum. The number of local minima outside that band diminishes exponentially with the size of the network. We empirically verify that the mathematical model exhibits similar behavior as the computer simulations, despite the presence of high dependencies in real networks. We conjecture that both simulated annealing and SGD converge to the band of low critical points, and that all critical points found there are local minima of high quality measured by the test error. This emphasizes a major difference between large- and small-size networks where for the latter poor quality local minima have non-zero probability of being recovered. Finally, we prove that recovering the global minimum becomes harder as the network size increases and that it is in practice irrelevant as global minimum often leads to overfitting.
References in corpus (1)
Cited by in corpus (108)
- Understanding deep learning requires rethinking generalization
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Non-convex Optimization for Machine Learning
- Spectral Norm Regularization for Improving the Generalizability of Deep Learning
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- MLPerf Training Benchmark
- Optimization for deep learning: theory and algorithms
- Overfitting Mechanism and Avoidance in Deep Neural Networks
- Mathematics of Deep Learning
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- The Power of Normalization: Faster Evasion of Saddle Points
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- The Landscape of Empirical Risk for Non-convex Losses
- Adversarial Examples in Modern Machine Learning: A Review
- Stochastic Cubic Regularization for Fast Nonconvex Optimization
- BlackMarks: Blackbox Multibit Watermarking for Deep Neural Networks
- Local minima in training of neural networks
- Gossip training for deep learning
- Convolutional Neural Networks Analyzed via Convolutional Sparse Coding
- A Multiple Filter Based Neural Network Approach to the Extrapolation of Adsorption Energies on Metal Surfaces for Catalysis Applications
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Coulomb GANs: Provably Optimal Nash Equilibria via Potential Fields
- Characterization of Gradient Dominance and Regularity Conditions for Neural Networks
- Convergent Block Coordinate Descent for Training Tikhonov Regularized Deep Neural Networks
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Negative eigenvalues of the Hessian in deep neural networks
- Recent advances in deep learning theory
- Loss Landscapes of Regularized Linear Autoencoders
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Stragglers Are Not Disaster: A Hybrid Federated Learning Algorithm with Delayed Gradients
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Large Scale Structure of Neural Network Loss Landscapes
- Generalization Bounds for Convolutional Neural Networks
- Quantitative Propagation of Chaos for SGD in Wide Neural Networks
- From Dependence to Causation
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- The number of saddles of the spherical -spin model
- A Correspondence Between Random Neural Networks and Statistical Field Theory
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
- Inference in Deep Networks in High Dimensions
- On the Stability of Deep Networks
- Exploring generative atomic models in cryo-EM reconstruction
- The Local Elasticity of Neural Networks
- Deep learning for time series classification
- Dynamic Inference with Neural Interpreters
- Learning-based Traffic State Reconstruction using Probe Vehicles
- Escaping Saddle Points Faster with Stochastic Momentum
- On Connected Sublevel Sets in Deep Learning
- Trusted Artificial Intelligence: Towards Certification of Machine Learning Applications
- Efficiently avoiding saddle points with zero order methods: No gradients required
- Saving Gradient and Negative Curvature Computations: Finding Local Minima More Efficiently
- Hessian Eigenspectra of More Realistic Nonlinear Models
- Neural Network Memorization Dissection
- The asymptotic spectrum of the Hessian of DNN throughout training
- Catch-Up Mix: Catch-Up Class for Struggling Filters in CNN
- Bregman Proximal Framework for Deep Linear Neural Networks
- Low-rank Bilinear Pooling for Fine-Grained Classification
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Supervised Deep Neural Networks (DNNs) for Pricing/Calibration of Vanilla/Exotic Options Under Various Different Processes
- Edge of chaos as a guiding principle for modern neural network training
- Tangent Space Separability in Feedforward Neural Networks
- Third-order Smoothness Helps: Even Faster Stochastic Optimization Algorithms for Finding Local Minima
- LV-ROVER: Lexicon Verified Recognizer Output Voting Error Reduction
- A Generative Model for Sampling High-Performance and Diverse Weights for Neural Networks
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- A Note on Connectivity of Sublevel Sets in Deep Learning
- WaveQ: Gradient-Based Deep Quantization of Neural Networks through Sinusoidal Adaptive Regularization
- A relativistic extension of Hopfield neural networks via the mechanical analogy
- Subaging in underparametrized Deep Neural Networks
- On the Convex Behavior of Deep Neural Networks in Relation to the Layers' Width
- Global Capacity Measures for Deep ReLU Networks via Path Sampling
- Induction, Popper, and machine learning
- Are Saddles Good Enough for Deep Learning?
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- High Efficient Reconstruction of Single-shot T2 Mapping from OverLapping-Echo Detachment Planar Imaging Based on Deep Residual Network
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- NeuSE: A Neural Snapshot Ensemble Method for Collaborative Filtering
- Search Spaces for Neural Model Training
- Largest Eigenvalues of the Conjugate Kernel of Single-Layered Neural Networks
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- Investigating the interaction between gradient-only line searches and different activation functions
- Deep Online Learning with Stochastic Constraints
- Conceptual capacity and effective complexity of neural networks
- Optimizing Shallow Networks for Binary Classification
- Acoustic Model Optimization Based On Evolutionary Stochastic Gradient Descent with Anchors for Automatic Speech Recognition
- Achieving Small Test Error in Mildly Overparameterized Neural Networks
- Novel Uncertainty Framework for Deep Learning Ensembles
- Some New Results for Poisson Binomial Models
- Universality of Gradient Descent Neural Network Training
- The Restricted Isometry of ReLU Networks: Generalization through Norm Concentration
- Weighted Aggregating Stochastic Gradient Descent for Parallel Deep Learning
- Large-Dimensional Random Matrix Theory and Its Applications in Deep Learning and Wireless Communications
- When Can Neural Networks Learn Connected Decision Regions?
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- On the Differentially Private Nature of Perturbed Gradient Descent
- BN-invariant sharpness regularizes the training model to better generalization
- A Comprehensive Study on Optimization Strategies for Gradient Descent In Deep Learning
- Learning Graph Neural Networks with Approximate Gradient Descent
- Algebraically-Informed Deep Networks (AIDN): A Deep Learning Approach to Represent Algebraic Structures
- Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian
- On the Second-order Convergence Properties of Random Search Methods
- Immunization of Pruning Attack in DNN Watermarking Using Constant Weight Code
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Solving hybrid machine learning tasks by traversing weight space geodesics
- How regularization affects the critical points in linear networks
- Escaping Saddle Points with Compressed SGD