Non-Gaussian processes and neural networks at finite widths
arXiv:1910.00019
Abstract
Gaussian processes are ubiquitous in nature and engineering. A case in point is a class of neural networks in the infinite-width limit, whose priors correspond to Gaussian processes. Here we perturbatively extend this correspondence to finite-width neural networks, yielding non-Gaussian processes as priors. The methodology developed herein allows us to track the flow of preactivation distributions by progressively integrating out random variables from lower to higher layers, reminiscent of renormalization-group flow. We further develop a perturbative procedure to perform Bayesian inference with weakly non-Gaussian priors.
33 pages, 3 figures; v2: final version accepted at MSML 2020, with some clarification on the connection to renormalization-group flow
References in corpus (5)
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Deep Neural Networks as Gaussian Processes
- Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
- An exact mapping between the Variational Renormalization Group and Deep Learning
- Asymptotics of Wide Networks from Feynman Diagrams
Cited by in corpus (8)
- Explaining Neural Scaling Laws
- Neural Networks and Quantum Field Theory
- Asymptotics of Wide Convolutional Neural Networks
- On the asymptotics of wide networks with polynomial activations
- Non-asymptotic approximations of neural networks by Gaussian processes
- Explicitly Bayesian Regularizations in Deep Learning
- Mixed Moments for the Product of Ginibre Matrices
- Conditional Deep Gaussian Processes: empirical Bayes hyperdata learning