Memory capacity of neural networks with threshold and ReLU activations
arXiv:2001.06938
Abstract
Overwhelming theoretical and empirical evidence shows that mildly overparametrized neural networks -- those with more connections than the size of the training data -- are often able to memorize the training data with accuracy. This was rigorously proved for networks with sigmoid activation functions and, very recently, for ReLU activations. Addressing a 1988 open question of Baum, we prove that this phenomenon holds for general multilayered perceptrons, i.e. neural networks with threshold activation functions, or with any mix of threshold and ReLU activations. Our construction is probabilistic and exploits sparsity.
26 pages. Minor inaccuracies corrected, discussion of prior work expanded
References in corpus (7)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Optimization for deep learning: theory and algorithms
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
- An Improved Analysis of Training Over-parameterized Deep Neural Networks
- Mildly Overparametrized Neural Nets can Memorize Training Data Efficiently
Cited by in corpus (8)
- Random Vector Functional Link Networks for Function Approximation on Manifolds
- Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory
- Approximation in shift-invariant spaces with deep ReLU neural networks
- Minimum Width for Universal Approximation
- Provable Memorization via Deep Neural Networks using Sub-linear Parameters
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- On the Optimal Memorization Power of ReLU Neural Networks
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology