Exact marginal prior distributions of finite Bayesian neural networks
arXiv:2104.11734
Abstract
Bayesian neural networks are theoretically well-understood only in the infinite-width limit, where Gaussian priors over network weights yield Gaussian priors over network outputs. Recent work has suggested that finite Bayesian networks may outperform their infinite counterparts, but their non-Gaussian function space priors have been characterized only though perturbative approaches. Here, we derive exact solutions for the function space priors for individual input examples of a class of finite fully-connected feedforward Bayesian neural networks. For deep linear networks, the prior has a simple expression in terms of the Meijer -function. The prior of a finite ReLU network is a mixture of the priors of linear networks of smaller widths, corresponding to different numbers of active units in each layer. Our results unify previous descriptions of finite network priors in terms of their tail decay and large-width behavior.
12+9 pages, 4 figures; v3: Accepted as NeurIPS 2021 Spotlight
References in corpus (6)
- What Are Bayesian Neural Network Posteriors Really Like?
- Bayesian Neural Network Priors Revisited
- A Correspondence Between Random Neural Networks and Statistical Field Theory
- On the asymptotics of wide networks with polynomial activations
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective
- Precise characterization of the prior predictive distribution of deep ReLU networks
Cited by in corpus (6)
- Depth induces scale-averaging in overparameterized linear Bayesian neural networks
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective
- Precise characterization of the prior predictive distribution of deep ReLU networks
- Bayesian neural network unit priors and generalized Weibull-tail property
- Dependence between Bayesian neural network units
- Conditional Deep Gaussian Processes: empirical Bayes hyperdata learning