Deep Neural Networks as Gaussian Processes
arXiv:1711.00165
Abstract
It has long been known that a single-layer fully-connected neural network with an i.i.d. prior over its parameters is equivalent to a Gaussian process (GP), in the limit of infinite network width. This correspondence enables exact Bayesian inference for infinite width neural networks on regression tasks by means of evaluating the corresponding GP. Recently, kernel functions which mimic multi-layer random neural networks have been developed, but only outside of a Bayesian framework. As such, previous work has not identified that these kernels can be used as covariance functions for GPs and allow fully Bayesian prediction with a deep neural network. In this work, we derive the exact equivalence between infinitely wide deep networks and GPs. We further develop a computationally efficient pipeline to compute the covariance function for these GPs. We then use the resulting GPs to perform Bayesian inference for wide deep neural networks on MNIST and CIFAR-10. We observe that trained neural network accuracy approaches that of the corresponding GP with increasing layer width, and that the GP uncertainty is strongly correlated with trained network prediction error. We further find that test performance increases as finite-width trained networks are made wider and more similar to a GP, and thus that GP predictions typically outperform those of finite-width networks. Finally we connect the performance of these GPs to the recent theory of signal propagation in random neural networks.
Published version in ICLR 2018. 10 pages + appendix
References in corpus (7)
- Adam: A Method for Stochastic Optimization
- Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
- Deep Gaussian Processes for Regression using Approximate Expectation Propagation
- Stochastic Gradient Descent as Approximate Bayesian Inference
- Steps Toward Deep Kernel Methods from Infinite Neural Networks
- Nested Variational Compression in Deep Gaussian Processes
- AutoGP: Exploring the Capabilities and Limitations of Gaussian Process Models
Cited by in corpus (210)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Machine learning and the physical sciences
- B-PINNs: Bayesian Physics-Informed Neural Networks for Forward and Inverse PDE Problems with Noisy Data
- On Exact Computation with an Infinitely Wide Neural Net
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- The Principles of Deep Learning Theory
- Multi-fidelity Bayesian Neural Networks: Algorithms and Applications
- Gradient Descent Finds Global Minima of Deep Neural Networks
- Scaling description of generalization with number of parameters in deep learning
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Explaining Neural Scaling Laws
- The Modern Mathematics of Deep Learning
- Multi-fidelity regression using artificial neural networks: efficient approximation of parameter-dependent output quantities
- When and why PINNs fail to train: A neural tangent kernel perspective
- On the Convergence Rate of Training Recurrent Neural Networks
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
- Deep Convolutional Networks as shallow Gaussian Processes
- Multi-fidelity surrogate modeling using long short-term memory networks
- Biology and medicine in the landscape of quantum advantages
- Neural Policy Gradient Methods: Global Optimality and Rates of Convergence
- Evaluating aleatoric and epistemic uncertainties of time series deep learning models for soil moisture predictions
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Representation Learning via Quantum Neural Tangent Kernels
- Neural-net-induced Gaussian process regression for function approximation and PDE solution
- Survey on Machine Learning for Traffic-Driven Service Provisioning in Optical Networks
- Neural Networks and Quantum Field Theory
- When Do Neural Networks Outperform Kernel Methods?
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- Learning Functional Priors and Posteriors from Data and Physics
- Insights on representational similarity in neural networks with canonical correlation
- Linear Multiple Low-Rank Kernel Based Stationary Gaussian Processes Regression for Time Series
- A Mean Field Theory of Batch Normalization
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- The Gaussian equivalence of generative models for learning with shallow neural networks
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- On the Selection of Initialization and Activation Function for Deep Neural Networks
- On the Expressiveness of Approximate Inference in Bayesian Neural Networks
- Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes
- A Fine-Grained Spectral Perspective on Neural Networks
- Physics-Driven Learning for Inverse Problems in Quantum Chromodynamics
- Nonperturbative renormalization for the neural network-QFT correspondence
- Sub-Weibull distributions: generalizing sub-Gaussian and sub-Exponential properties to heavier-tailed distributions
- Variational Bayes survival analysis for unemployment modelling
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? -- A Neural Tangent Kernel Perspective
- 'In-Between' Uncertainty in Bayesian Neural Networks
- Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed
- Bayesian Deep Ensembles via the Neural Tangent Kernel
- Learning through atypical "phase transitions" in overparameterized neural networks
- Mathematical Models of Overparameterized Neural Networks
- Asymptotics of representation learning in finite Bayesian neural networks
- Bayesian Neural Network Priors Revisited
- Data-driven deep learning algorithms for time-varying infection rates of COVID-19 and mitigation measures
- Data-driven Aerodynamic Analysis of Structures using Gaussian Processes
- Improving Local Identifiability in Probabilistic Box Embeddings
- Exploring the Function Space of Deep-Learning Machines
- Adversarial Reprogramming of Neural Networks
- Dataset Distillation with Infinitely Wide Convolutional Networks
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Wide and Deep Neural Networks Achieve Optimality for Classification
- Bayesian Neural Networks for Fast SUSY Predictions
- Some models are useful, but how do we know which ones? Towards a unified Bayesian model taxonomy
- Finite size corrections for neural network Gaussian processes
- Algorithmic Linearly Constrained Gaussian Processes
- Tensor Programs III: Neural Matrix Laws
- Deeper Connections between Neural Networks and Gaussian Processes Speed-up Active Learning
- Gaussian Process States: A data-driven representation of quantum many-body physics
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networks
- Repulsive Deep Ensembles are Bayesian
- Depth induces scale-averaging in overparameterized linear Bayesian neural networks
- Non-parametric Bayesian approach to extrapolation problems in configuration interaction methods
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
- Locality defeats the curse of dimensionality in convolutional teacher-student scenarios
- The Benefits of Over-parameterization at Initialization in Deep ReLU Networks
- FedLoc: Federated Learning Framework for Data-Driven Cooperative Localization and Location Data Processing
- Towards Quantification of Assurance for Learning-enabled Components
- Robust Pruning at Initialization
- On the infinite width limit of neural networks with a standard parameterization
- On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization
- Non-Gaussian processes and neural networks at finite widths
- A Gaussian Process perspective on Convolutional Neural Networks
- On the Similarity between the Laplace and Neural Tangent Kernels
- NN2Poly: A polynomial representation for deep feed-forward artificial neural networks
- Quantifying Point-Prediction Uncertainty in Neural Networks via Residual Estimation with an I/O Kernel
- Approximate Inference Turns Deep Networks into Gaussian Processes
- Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks
- Effect of Activation Functions on the Training of Overparametrized Neural Nets
- Robust Bayesian Target Value Optimization
- Deep Learning in Wide-field Surveys: Fast Analysis of Strong Lenses in Ground-based Cosmic Experiments
- Defence against adversarial attacks using classical and quantum-enhanced Boltzmann machines
- Neural networks are a priori biased towards Boolean functions with low entropy
- Order and Chaos: NTK views on DNN Normalization, Checkerboard and Boundary Artifacts
- On the expected behaviour of noise regularised deep neural networks as Gaussian processes
- Exact marginal prior distributions of finite Bayesian neural networks
- Laziness, Barren Plateau, and Noise in Machine Learning
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime
- Convex Geometry and Duality of Over-parameterized Neural Networks
- Quantum Machine Learning using Gaussian Processes with Performant Quantum Kernels
- Variational Implicit Processes
- Deep Equals Shallow for ReLU Networks in Kernel Regimes
- Pathological spectra of the Fisher information metric and its variants in deep neural networks
- Is SGD a Bayesian sampler? Well, almost
- Stationary Activations for Uncertainty Calibration in Deep Learning
- Infinite Neural Network Quantum States: Entanglement and Training Dynamics
- On the asymptotics of wide networks with polynomial activations
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- A Batched Scalable Multi-Objective Bayesian Optimization Algorithm
- Adversarial Robustness Guarantees for Classification with Gaussian Processes
- Space of Functions Computed by Deep-Layered Machines
- Quantum-enhanced neural networks in the neural tangent kernel framework
- Artificial Neural Network Modeling for Airline Disruption Management
- Gating creates slow modes and controls phase-space complexity in GRUs and LSTMs
- On neural network kernels and the storage capacity problem
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Implicit Rugosity Regularization via Data Augmentation
- Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective
- On Infinite-Width Hypernetworks
- Information Geometry of Orthogonal Initializations and Training
- How Wrong Am I? - Studying Adversarial Examples and their Impact on Uncertainty in Gaussian Process Machine Learning Models
- Wide Neural Networks with Bottlenecks are Deep Gaussian Processes
- Predicting Training Time Without Training
- The Recurrent Neural Tangent Kernel
- The Ridgelet Prior: A Covariance Function Approach to Prior Specification for Bayesian Neural Networks
- Initialization of ReLUs for Dynamical Isometry
- Exact posterior distributions of wide Bayesian neural networks
- Gaussian Process-Gated Hierarchical Mixtures of Experts
- A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off
- Approximation and Learning with Deep Convolutional Models: a Kernel Perspective
- Mitigating Uncertainty in Document Classification
- Why bigger is not always better: on finite and infinite neural networks
- Dimensional Reweighting Graph Convolutional Networks
- Deep Feature Gaussian Processes for Single-Scene Aerosol Optical Depth Reconstruction
- Correlated Weights in Infinite Limits of Deep Convolutional Neural Networks
- Gradient descent in Gaussian random fields as a toy model for high-dimensional optimisation in deep learning
- Bayesian Image Classification with Deep Convolutional Gaussian Processes
- Bayesian Learning of LF-MMI Trained Time Delay Neural Networks for Speech Recognition
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- On Random Kernels of Residual Architectures
- Infinitely Wide Tensor Networks as Gaussian Process
- Accelerating the training of single-layer binary neural networks using the HHL quantum algorithm
- A brief note on understanding neural networks as Gaussian processes
- Kernel-Based Smoothness Analysis of Residual Networks
- Real-time regression analysis with deep convolutional neural networks
- Sparse Uncertainty Representation in Deep Learning with Inducing Weights
- Why flatness does and does not correlate with generalization for deep neural networks
- Scale Mixtures of Neural Network Gaussian Processes
- Epistemic Neural Networks
- Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width Limit
- Bayesian RG Flow in Neural Network Field Theories
- Are wider nets better given the same number of parameters?
- Adversarial Robustness Guarantees for Random Deep Neural Networks
- Infinite-channel deep stable convolutional neural networks
- An Infinite-Feature Extension for Bayesian ReLU Nets That Fixes Their Asymptotic Overconfidence
- Wide stochastic networks: Gaussian limit and PAC-Bayesian training
- Mean field theory for deep dropout networks: digging up gradient backpropagation deeply
- Label-Aware Neural Tangent Kernel: Toward Better Generalization and Local Elasticity
- Trap of Feature Diversity in the Learning of MLPs
- Large-width functional asymptotics for deep Gaussian neural networks
- Scaling Neural Tangent Kernels via Sketching and Random Features
- Uncertainty Quantification From Scaling Laws in Deep Neural Networks
- Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory?
- Uniform Priors for Data-Efficient Transfer
- Neural Network Gaussian Process Considering Input Uncertainty for Composite Structures Assembly
- Deep kernel processes
- Initializing ReLU networks in an expressive subspace of weights
- Scalable neural network-based blackbox optimization
- A Dynamical Central Limit Theorem for Shallow Neural Networks
- Posterior contraction rates for constrained deep Gaussian processes in density estimation and classication
- Learning curves for Gaussian process regression with power-law priors and targets
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Dependence between Bayesian neural network units
- Predicting the outputs of finite deep neural networks trained with noisy gradients
- The Bayesian Method of Tensor Networks
- Deep kernel learning for integral measurements
- Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling
- Out-of-Distribution Generalization in Kernel Regression
- Uniform Convergence, Adversarial Spheres and a Simple Remedy
- Dynamical transition in controllable quantum neural networks with large depth
- Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
- Location Trace Privacy Under Conditional Priors
- On the Preservation of Spatio-temporal Information in Machine Learning Applications
- Deep covariate-learning: optimising information extraction from terrain texture for geostatistical modelling applications
- Deep Latent-Variable Kernel Learning
- A Statistician Teaches Deep Learning
- Catapult Dynamics and Phase Transitions in Quadratic Nets
- WGAN with an Infinitely Wide Generator Has No Spurious Stationary Points
- Faster Kernel Interpolation for Gaussian Processes
- Enhanced Recurrent Neural Tangent Kernels for Non-Time-Series Data
- Learning with Neural Tangent Kernels in Near Input Sparsity Time
- Parametric machines: a fresh approach to architecture search
- Sparsity-Probe: Analysis tool for Deep Learning Models
- Likelihood-Free Gaussian Process for Regression
- Ghosts in Neural Networks: Existence, Structure and Role of Infinite-Dimensional Null Space
- A variational approximate posterior for the deep Wishart process
- Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks
- Implicit Bias of Linear Equivariant Networks
- New Insights into Graph Convolutional Networks using Neural Tangent Kernels
- Covariate Shift in High-Dimensional Random Feature Regression
- Discriminative Clustering with Representation Learning with any Ratio of Labeled to Unlabeled Data
- PAC-Bayesian Bounds for Deep Gaussian Processes
- On the relationship between multitask neural networks and multitask Gaussian Processes
- Simulation Example of a Black Noise
- Sequential online prediction in the presence of outliers and change points: an instant temporal structure learning approach
- A note on regularised NTK dynamics with an application to PAC-Bayesian training
- Input Modeling and Uncertainty Quantification for Improving Volatile Residential Load Forecasting
- Large-Scale Learning with Fourier Features and Tensor Decompositions
- Uncertainty-Aware Multi-Modal Ensembling for Severity Prediction of Alzheimer's Dementia
- Generalization Properties of hyper-RKHS and its Applications