Depth induces scale-averaging in overparameterized linear Bayesian neural networks
arXiv:2111.11954 · doi:10.1109/IEEECONF53345.2021.9723137
Abstract
Inference in deep Bayesian neural networks is only fully understood in the infinite-width limit, where the posterior flexibility afforded by increased depth washes out and the posterior predictive collapses to a shallow Gaussian process. Here, we interpret finite deep linear Bayesian neural networks as data-dependent scale mixtures of Gaussian process predictors across output channels. We leverage this observation to study representation learning in these networks, allowing us to connect limiting results obtained in previous studies within a unified framework. In total, these results advance our analytical understanding of how depth affects inference in a simple class of Bayesian neural networks.
8 pages, no figures
References in corpus (10)
- The Principles of Deep Learning Theory
- Statistical Mechanics of Deep Linear Neural Networks: The Back-Propagating Kernel Renormalization
- Finite Versus Infinite Neural Networks: an Empirical Study
- Statistical limits of dictionary learning: random matrix theory and the spectral replica method
- Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising
- Asymptotics of representation learning in finite Bayesian neural networks
- A Unifying View on Implicit Bias in Training Linear Neural Networks
- Exact marginal prior distributions of finite Bayesian neural networks
- Scale Mixtures of Neural Network Gaussian Processes
- Neural Networks as Kernel Learners: The Silent Alignment Effect
Cited by in corpus (4)
- A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit
- Asymptotics of representation learning in finite Bayesian neural networks
- Bayesian RG Flow in Neural Network Field Theories
- Summary statistics of learning link changing neural representations to behavior