Repulsive Deep Ensembles are Bayesian
arXiv:2106.11642
Abstract
Deep ensembles have recently gained popularity in the deep learning community for their conceptual simplicity and efficiency. However, maintaining functional diversity between ensemble members that are independently trained with gradient descent is challenging. This can lead to pathologies when adding more ensemble members, such as a saturation of the ensemble performance, which converges to the performance of a single model. Moreover, this does not only affect the quality of its predictions, but even more so the uncertainty estimates of the ensemble, and thus its performance on out-of-distribution data. We hypothesize that this limitation can be overcome by discouraging different ensemble members from collapsing to the same function. To this end, we introduce a kernelized repulsive term in the update rule of the deep ensembles. We show that this simple modification not only enforces and maintains diversity among the members but, even more importantly, transforms the maximum a posteriori inference into proper Bayesian inference. Namely, we show that the training dynamics of our proposed repulsive ensembles follow a Wasserstein gradient flow of the KL divergence with the true posterior. We study repulsive terms in weight and function space and empirically compare their performance to standard ensembles and Bayesian baselines on synthetic and real-world prediction tasks.
References in corpus (14)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Weight Uncertainty in Neural Networks
- Understanding deep learning requires rethinking generalization
- Functional Variational Bayesian Neural Networks
- What Are Bayesian Neural Network Posteriors Really Like?
- Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
- Bayesian Neural Network Priors Revisited
- BNNpriors: A library for Bayesian neural network inference with different prior distributions
- On Linear Identifiability of Learned Representations
- Annealed Stein Variational Gradient Descent
- Understanding Variational Inference in Function-Space
- A Bayesian Perspective on Training Speed and Model Selection
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- Posterior Meta-Replay for Continual Learning
Cited by in corpus (8)
- Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
- Bayesian posterior approximation with stochastic ensembles
- Quantum Bayesian Neural Networks
- Deep Classifiers with Label Noise Modeling and Distance Awareness
- Pathologies in priors and inference for Bayesian transformers
- Greedy Bayesian Posterior Approximation with Deep Ensembles
- Neural Variational Gradient Descent
- A Bayesian Approach to Invariant Deep Neural Networks