Asymptotics of Wide Networks from Feynman Diagrams
arXiv:1909.11304
Abstract
Understanding the asymptotic behavior of wide networks is of considerable interest. In this work, we present a general method for analyzing this large width behavior. The method is an adaptation of Feynman diagrams, a standard tool for computing multivariate Gaussian integrals. We apply our method to study training dynamics, improving existing bounds and deriving new results on wide network evolution during stochastic gradient descent. Going beyond the strict large width limit, we present closed-form expressions for higher-order terms governing wide network training, and test these predictions empirically.
10 pages, 3 figures, 1 Table + Appendices
References in corpus (6)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Gradient Descent Happens in a Tiny Subspace
- Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- A Correspondence Between Random Neural Networks and Statistical Field Theory
- Limitations of Lazy Training of Two-layers Neural Networks
Cited by in corpus (11)
- Representation Learning via Quantum Neural Tangent Kernels
- Nonperturbative renormalization for the neural network-QFT correspondence
- Asymptotics of representation learning in finite Bayesian neural networks
- Laziness, Barren Plateau, and Noise in Machine Learning
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- On the asymptotics of wide networks with polynomial activations
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Bayesian RG Flow in Neural Network Field Theories
- Dynamical transition in controllable quantum neural networks with large depth
- Understanding Deflation Process in Over-parametrized Tensor Decomposition
- Catapult Dynamics and Phase Transitions in Quadratic Nets