On Connected Sublevel Sets in Deep Learning
arXiv:1901.07417
Abstract
This paper shows that every sublevel set of the loss function of a class of deep over-parameterized neural nets with piecewise linear activation functions is connected and unbounded. This implies that the loss has no bad local valleys and all of its global minima are connected within a unique and potentially very large global valley.
Accepted at ICML 2019. More discussions and visualizations are added
References in corpus (7)
- The Loss Surfaces of Multilayer Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- On the Computational Efficiency of Training Neural Networks
- Diverse Neural Network Learns True Target Functions
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Eliminating all bad Local Minima from Loss Landscapes without even adding an Extra Unit