The Principles of Deep Learning Theory
arXiv:2106.10165 · doi:10.1017/9781009023405
Abstract
This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks, we explain how to determine an accurate description of the output of trained networks by solving layer-to-layer iteration equations and nonlinear learning dynamics. A main result is that the predictions of networks are described by nearly-Gaussian distributions, with the depth-to-width aspect ratio of the network controlling the deviations from the infinite-width Gaussian description. We explain how these effectively-deep networks learn nontrivial representations from training and more broadly analyze the mechanism of representation learning for nonlinear models. From a nearly-kernel-methods perspective, we find that the dependence of such models' predictions on the underlying learning algorithm can be expressed in a simple and universal way. To obtain these results, we develop the notion of representation group flow (RG flow) to characterize the propagation of signals through the network. By tuning networks to criticality, we give a practical solution to the exploding and vanishing gradient problem. We further explain how RG flow leads to near-universal behavior and lets us categorize networks built from different activation functions into universality classes. Altogether, we show that the depth-to-width ratio governs the effective model complexity of the ensemble of trained networks. By using information-theoretic techniques, we estimate the optimal aspect ratio at which we expect the network to be practically most useful and show how residual connections can be used to push this scale to arbitrary depths. With these tools, we can learn in detail about the inductive bias of architectures, hyperparameters, and optimizers.
471 pages, to be published by Cambridge University Press; v2: hyperlinks fixed, index added
References in corpus (5)
Cited by in corpus (38)
- Analytic theory for the dynamics of wide quantum neural networks
- Dimension of activity in random neural networks
- Nonperturbative renormalization for the neural network-QFT correspondence
- A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit
- A perspective on machine learning and data science for strongly correlated electron problems
- Asymptotics of representation learning in finite Bayesian neural networks
- Contrasting random and learned features in deep Bayesian linear regression
- The Quantum Path Kernel: a Generalized Quantum Neural Tangent Kernel for Deep Quantum Machine Learning
- Gell-Mann-Low criticality in neural networks
- Depth induces scale-averaging in overparameterized linear Bayesian neural networks
- The Gaia-ESO Survey: Chemical evolution of Mg and Al in the Milky Way with Machine-Learning
- Prediction of Transportation Index for Urban Patterns in Small and Medium-sized Indian Cities using Hybrid RidgeGAN Model
- Decomposing neural networks as mappings of correlation functions
- Differentiable Physics: A Position Piece
- Laziness, Barren Plateau, and Noise in Machine Learning
- Amplitude-assisted tagging of longitudinally polarised bosons using wide neural networks
- infomeasure: A Comprehensive Python Package for Information Theory Measures and Estimators
- Coding schemes in neural networks learning classification tasks
- Wavelet Conditional Renormalization Group
- Emergence of hierarchical modes from deep learning
- Infinite Neural Network Quantum States: Entanglement and Training Dynamics
- Thermodynamics of the Ising model encoded in restricted Boltzmann machines
- Precise characterization of the prior predictive distribution of deep ReLU networks
- Machine-Learning Performance on Higgs-Pair Production Associated with Dark Matter at the LHC
- Ambitions for theory in the physics of life
- Neural Tangent Kernel Maximum Mean Discrepancy
- Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
- Bayesian RG Flow in Neural Network Field Theories
- Weight fluctuations in (deep) linear neural networks and a derivation of the inverse-variance flatness relation
- Large and moderate deviations for Gaussian neural networks
- Quantitative convergence of trained quantum neural networks to a Gaussian process
- Appearance of Random Matrix Theory in Deep Learning
- Dynamical transition in controllable quantum neural networks with large depth
- Machine learning that predicts well may not learn the correct physical descriptions of glassy systems
- Dynamic neuron approach to deep neural networks: Decoupling neurons for renormalization group analysis
- Gauge-covariant stochastic neural fields: Stability and finite-width effects
- Bayesian inference with finitely wide neural networks
- Statistics of correlations in nonlinear recurrent neural networks