paper

Bayesian Inference with Deep Weakly Nonlinear Networks

arXiv:2405.16630

Abstract

We show at a physics level of rigor that Bayesian inference with a fully connected neural network and a shaped nonlinearity of the form is (perturbatively) solvable in the regime where the number of training datapoints , the input dimension , the network layer widths , and the network depth are simultaneously large. Our results hold with weak assumptions on the data; the main constraint is that . We provide techniques to compute the model evidence and posterior to arbitrary order in and at arbitrary temperature. We report the following results from the first-order computation: 1. When the width is much larger than the depth and training set size , neural network Bayesian inference coincides with Bayesian inference using a kernel. The value of determines the curvature of a sphere, hyperbola, or plane into which the training data is implicitly embedded under the feature map. 2. When is a small constant, neural network Bayesian inference departs from the kernel regime. At zero temperature, neural network Bayesian inference is equivalent to Bayesian inference using a data-dependent kernel, and serves as an effective depth that controls the extent of feature learning. 3. In the restricted case of deep linear networks () and noisy data, we show a simple data model for which evidence and generalization error are optimal at zero temperature. As increases, both evidence and generalization further improve, demonstrating the benefit of depth in benign overfitting.