The Multilinear Structure of ReLU Networks
arXiv:1712.10132
Abstract
We study the loss surface of neural networks equipped with a hinge loss criterion and ReLU or leaky ReLU nonlinearities. Any such network defines a piecewise multilinear form in parameter space. By appealing to harmonic analysis we show that all local minima of such network are non-differentiable, except for those minima that occur in a region of parameter space where the loss surface is perfectly flat. Non-differentiable minima are therefore not technicalities or pathologies; they are heart of the problem when investigating the loss of ReLU networks. As a consequence, we must employ techniques from nonsmooth analysis to study these loss surfaces. We show how to apply these techniques in some illustrative cases.
References in corpus (6)
- The Loss Surfaces of Multilayer Networks
- Recovery Guarantees for One-hidden-layer Neural Networks
- The Landscape of Empirical Risk for Non-convex Losses
- When is a Convolutional Filter Easy To Learn?
- SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data
- Deep linear neural networks with arbitrary loss: All local minima are global
Cited by in corpus (9)
- Optimization for deep learning: theory and algorithms
- Point Spread Function Modelling for Wide Field Small Aperture Telescopes with a Denoising Autoencoder
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- Learning ReLU Networks via Alternating Minimization
- Are deep ResNets provably better than linear predictors?
- The Effects of Mild Over-parameterization on the Optimization Landscape of Shallow ReLU Neural Networks
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- Achieving Small Test Error in Mildly Overparameterized Neural Networks