Optimization Landscape and Expressivity of Deep CNNs
arXiv:1710.10928
Abstract
We analyze the loss landscape and expressiveness of practical deep convolutional neural networks (CNNs) with shared weights and max pooling layers. We show that such CNNs produce linearly independent features at a "wide" layer which has more neurons than the number of training samples. This condition holds e.g. for the VGG network. Furthermore, we provide for such wide CNNs necessary and sufficient conditions for global minima with zero training error. For the case where the wide layer is followed by a fully connected layer we show that almost every critical point of the empirical loss is a global minimum with zero training error. Our analysis suggests that both depth and width are very important in deep learning. While depth brings more representational power and allows the network to learn high level features, width smoothes the optimization landscape of the loss function in the sense that a sufficiently wide network has a well-behaved loss surface with almost no bad local minima.
Accepted at ICML 2018
Cited by in corpus (31)
- Hyperspectral Image Classification-Traditional to Deep Models: A Survey for Future Prospects
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- Are All Layers Created Equal?
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Effect of Depth and Width on Local Minima in Deep Learning
- Elimination of All Bad Local Minima in Deep Learning
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions
- Robust Pruning at Initialization
- Visualized Insights into the Optimization Landscape of Fully Convolutional Networks
- Provable Memorization via Deep Neural Networks using Sub-linear Parameters
- On Connected Sublevel Sets in Deep Learning
- Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- Non-attracting Regions of Local Minima in Deep and Wide Neural Networks
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural Networks
- Increasing Depth Leads to U-Shaped Test Risk in Over-parameterized Convolutional Networks
- Information-Theoretic Local Minima Characterization and Regularization
- A Note on Connectivity of Sublevel Sets in Deep Learning
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
- On the Optimal Memorization Power of ReLU Neural Networks
- Optimization Landscapes of Wide Deep Neural Networks Are Benign
- On the Global Convergence of Gradient Descent for multi-layer ResNets in the mean-field regime
- Measure, Manifold, Learning, and Optimization: A Theory Of Neural Networks
- Theoretical Exploration of Flexible Transmitter Model
- A Generative Neural Network Framework for Automated Software Testing
- When Can Neural Networks Learn Connected Decision Regions?
- Benefits of over-parameterization with EM
- When Are Solutions Connected in Deep Networks?