An analytic theory of shallow networks dynamics for hinge loss classification
arXiv:2006.11209 · doi:10.1088/1742-5468/ac3a76
Abstract
Neural networks have been shown to perform incredibly well in classification tasks over structured high-dimensional datasets. However, the learning dynamics of such networks is still poorly understood. In this paper we study in detail the training dynamics of a simple type of neural network: a single hidden layer trained to perform a classification task. We show that in a suitable mean-field limit this case maps to a single-node learning problem with a time-dependent dataset determined self-consistently from the average nodes population. We specialize our theory to the prototypical case of a linearly separable dataset and a linear hinge loss, for which the dynamics can be explicitly solved. This allow us to address in a simple setting several phenomena appearing in modern networks such as slowing down of training dynamics, crossover between rich and lazy learning, and overfitting. Finally, we asses the limitations of mean-field theory by studying the case of large but finite number of nodes and of training samples.
16 pages, 6 figures
References in corpus (10)
- On Exact Computation with an Infinitely Wide Neural Net
- On Lazy Training in Differentiable Programming
- Scaling description of generalization with number of parameters in deep learning
- Kernel and Rich Regimes in Overparametrized Models
- An analytic theory of generalization dynamics and transfer learning in deep linear networks
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- A mean-field limit for certain deep neural networks
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks
- Generalisation dynamics of online learning in over-parameterised neural networks