Bayesian Deep Learning and a Probabilistic Perspective of Generalization
arXiv:2002.08791
Abstract
The key distinguishing property of a Bayesian approach is marginalization, rather than using a single setting of weights. Bayesian marginalization can particularly improve the accuracy and calibration of modern deep neural networks, which are typically underspecified by the data, and can represent many compelling but different solutions. We show that deep ensembles provide an effective mechanism for approximate Bayesian marginalization, and propose a related approach that further improves the predictive distribution by marginalizing within basins of attraction, without significant overhead. We also investigate the prior over functions implied by a vague distribution over neural network weights, explaining the generalization properties of such models from a probabilistic perspective. From this perspective, we explain results that have been presented as mysterious and distinct to neural network generalization, such as the ability to fit images with random labels, and show that these results can be reproduced with Gaussian processes. We also show that Bayesian model averaging alleviates double descent, resulting in monotonic performance improvements with increased flexibility. Finally, we provide a Bayesian perspective on tempering for calibrating predictive distributions.
31 pages, 19 figures
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- On Calibration of Modern Neural Networks
- Expectation Propagation for approximate Bayesian inference
- Weight Uncertainty in Neural Networks
- Understanding deep learning requires rethinking generalization
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Deep Ensembles: A Loss Landscape Perspective
- Fantastic Generalization Measures and Where to Find Them
- Functional Variational Bayesian Neural Networks
- The Case for Bayesian Deep Learning
- Optimal Regularization Can Mitigate Double Descent
- Introducing an Explicit Symplectic Integration Scheme for Riemannian Manifold Hamiltonian Monte Carlo
- Output-Constrained Bayesian Neural Networks
Cited by in corpus (61)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Hands-on Bayesian Neural Networks -- a Tutorial for Deep Learning Users
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- What Are Bayesian Neural Network Posteriors Really Like?
- Getting a CLUE: A Method for Explaining Uncertainty Estimates
- Finite Versus Infinite Neural Networks: an Empirical Study
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Hyperparameter Ensembles for Robustness and Uncertainty Quantification
- Uncertainty Quantification and Deep Ensembles
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- Rethinking Parameter Counting in Deep Models: Effective Dimensionality Revisited
- Bayesian Deep Ensembles via the Neural Tangent Kernel
- Asymptotics of representation learning in finite Bayesian neural networks
- A Comparison of Uncertainty Estimation Approaches in Deep Learning Components for Autonomous Vehicle Applications
- BNNpriors: A library for Bayesian neural network inference with different prior distributions
- Neural Ensemble Search for Uncertainty Estimation and Dataset Shift
- On Power Laws in Deep Ensembles
- Repulsive Deep Ensembles are Bayesian
- Learning under Model Misspecification: Applications to Variational and Ensemble methods
- A Quantitative Comparison of Epistemic Uncertainty Maps Applied to Multi-Class Segmentation
- Generalization bounds for deep learning
- A Simple Probabilistic Method for Deep Classification under Input-Dependent Label Noise
- Bayesian active learning for production, a systematic study and a reusable library
- Exact marginal prior distributions of finite Bayesian neural networks
- Dimensionality reduction, regularization, and generalization in overparameterized regressions
- Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints
- A Bayesian Perspective on Training Speed and Model Selection
- Scalable marginalization of correlated latent variables with applications to learning particle interaction kernels
- Is SGD a Bayesian sampler? Well, almost
- Deep transformation models for functional outcome prediction after acute ischemic stroke
- Dangers of Bayesian Model Averaging under Covariate Shift
- Deep Ensembles from a Bayesian Perspective
- LiBRe: A Practical Bayesian Approach to Adversarial Detection
- Precise characterization of the prior predictive distribution of deep ReLU networks
- On Stein Variational Neural Network Ensembles
- Simulation-Based Inference with Approximately Correct Parameters via Maximum Entropy
- A benchmark study on reliable molecular supervised learning via Bayesian learning
- Training on Test Data with Bayesian Adaptation for Covariate Shift
- A comprehensive study on the prediction reliability of graph neural networks for virtual screening
- Structured Weight Priors for Convolutional Neural Networks
- A Bayesian neural network predicts the dissolution of compact planetary systems
- Disentangling the Roles of Curation, Data-Augmentation and the Prior in the Cold Posterior Effect
- Structured Dropout Variational Inference for Bayesian Neural Networks
- An information-based metric for observing strategy optimization, demonstrated in the context of photometric redshifts with applications to cosmology
- Variational Auto-Regressive Gaussian Processes for Continual Learning
- Noisy Training Improves E2E ASR for the Edge
- Distributional Gaussian Process Layers for Outlier Detection in Image Segmentation
- Quantifying Uncertainty in Deep Spatiotemporal Forecasting
- Intrinsic uncertainties and where to find them
- Calibration and Uncertainty Quantification of Bayesian Convolutional Neural Networks for Geophysical Applications
- Who's Afraid of Thomas Bayes?
- Posterior Temperature Optimization in Variational Inference for Inverse Problems
- Identifying and Exploiting Structures for Reliable Deep Learning
- Contributions to Large Scale Bayesian Inference and Adversarial Machine Learning
- Pathologies in priors and inference for Bayesian transformers
- Self-Reflective Variational Autoencoder
- The Effect of Prior Lipschitz Continuity on the Adversarial Robustness of Bayesian Neural Networks
- Improving Uncertainty Calibration via Prior Augmented Data
- Greedy Bayesian Posterior Approximation with Deep Ensembles
- Rapid Risk Minimization with Bayesian Models Through Deep Learning Approximation