Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning
arXiv:2002.06470
Abstract
Uncertainty estimation and ensembling methods go hand-in-hand. Uncertainty estimation is one of the main benchmarks for assessment of ensembling performance. At the same time, deep learning ensembles have provided state-of-the-art results in uncertainty estimation. In this work, we focus on in-domain uncertainty for image classification. We explore the standards for its quantification and point out pitfalls of existing metrics. Avoiding these pitfalls, we perform a broad study of different ensembling techniques. To provide more insight in this study, we introduce the deep ensemble equivalent score (DEE) and show that many sophisticated ensembling techniques are equivalent to an ensemble of only few independently trained networks in terms of test performance.
Cited by in corpus (43)
- Machine Learning and Deep Learning -- A review for Ecologists
- A Review and Comparative Study on Probabilistic Object Detection in Autonomous Driving
- What Are Bayesian Neural Network Posteriors Really Like?
- Deeply Uncertain: Comparing Methods of Uncertainty Quantification in Deep Learning Algorithms
- i-PI 3.0: a flexible and efficient framework for advanced atomistic simulations
- Greedy Policy Search: A Simple Baseline for Learnable Test-Time Augmentation
- Uncertainty in Gradient Boosting via Ensembles
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Uncertainty Estimation in Autoregressive Structured Prediction
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- All You Need is a Good Functional Prior for Bayesian Deep Learning
- Twin Neural Network Regression
- Uncertainty-Aware Machine Translation Evaluation
- Regression Prior Networks
- Neural Ensemble Search for Uncertainty Estimation and Dataset Shift
- On Power Laws in Deep Ensembles
- Diversity Matters When Learning From Ensembles
- Bayesian Deep Learning via Subnetwork Inference
- A Unified Benchmark for the Unknown Detection Capability of Deep Neural Networks
- Combining Ensembles and Data Augmentation can Harm your Calibration
- Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints
- Estimating and Evaluating Regression Predictive Uncertainty in Deep Object Detectors
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- Robust Semantic Segmentation with Superpixel-Mix
- NOMU: Neural Optimization-based Model Uncertainty
- Mean-Field Approximation to Gaussian-Softmax Integral with Application to Uncertainty Estimation
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior Perspective
- Deep Ensembles on a Fixed Memory Budget: One Wide Network or Several Thinner Ones?
- Effect of the output activation function on the probabilities and errors in medical image segmentation
- An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
- Few-shot Conformal Prediction with Auxiliary Tasks
- Why have a Unified Predictive Uncertainty? Disentangling it using Deep Split Ensembles
- Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift
- Robustness via Cross-Domain Ensembles
- Efficient Conformal Prediction via Cascaded Inference with Expanded Admission
- Mixtures of Laplace Approximations for Improved Post-Hoc Uncertainty in Deep Learning
- Reliability Quantification of Deep Reinforcement Learning-based Control
- Revisiting Explicit Regularization in Neural Networks for Well-Calibrated Predictive Uncertainty
- Neural Bootstrapper
- Identifying and Exploiting Structures for Reliable Deep Learning
- Why Calibration Error is Wrong Given Model Uncertainty: Using Posterior Predictive Checks with Deep Learning
- Greedy Bayesian Posterior Approximation with Deep Ensembles
- Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy Performance