DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
arXiv:2101.05544
Abstract
Deep ensembles perform better than a single network thanks to the diversity among their members. Recent approaches regularize predictions to increase diversity; however, they also drastically decrease individual members' performances. In this paper, we argue that learning strategies for deep ensembles need to tackle the trade-off between ensemble diversity and individual accuracies. Motivated by arguments from information theory and leveraging recent advances in neural estimation of conditional mutual information, we introduce a novel training criterion called DICE: it increases diversity by reducing spurious correlations among features. The main idea is that features extracted from pairs of members should only share information useful for target class prediction without being conditionally redundant. Therefore, besides the classification loss with information bottleneck, we adversarially prevent features from being conditionally predictable from each other. We manage to reduce simultaneous errors while protecting class information. We obtain state-of-the-art accuracy results on CIFAR-10/100: for example, an ensemble of 5 networks trained with DICE matches an ensemble of 7 networks trained independently. We further analyze the consequences on calibration, uncertainty estimation, out-of-distribution detection and online co-distillation.
Published as a conference paper at ICLR 2021. 9 main pages, 13 figures, 12 tables
References in corpus (30)
- Auto-Encoding Variational Bayes
- Distilling the Knowledge in a Neural Network
- Estimating Mutual Information
- Learning deep representations by mutual information estimation and maximization
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
- What Makes for Good Views for Contrastive Learning?
- Contrastive Multiview Coding
- Learning Confidence for Out-of-Distribution Detection in Neural Networks
- TurkerGaze: Crowdsourcing Saliency with Webcam based Eye Tracking
- Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks
- Improving Adversarial Robustness via Promoting Ensemble Diversity
- Measuring Calibration in Deep Learning
- The Conditional Entropy Bottleneck
- Improving Adversarial Robustness of Ensembles with Diversity Training
- Feature-map-level Online Adversarial Knowledge Distillation
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning
- Uncertainty in the Variational Information Bottleneck
- Rethinking Softmax with Cross-Entropy: Neural Network Classifier as Mutual Information Estimator
- Feature Fusion for Online Mutual Knowledge Distillation
- EMI: Exploration with Mutual Information
- Visual Representations: Defining Properties and Deep Approximations
- An Unsupervised Information-Theoretic Perceptual Quality Metric
- Diverse Ensembles Improve Calibration
- Diversity inducing Information Bottleneck in Model Ensembles
- Unpacking Information Bottlenecks: Unifying Information-Theoretic Objectives in Deep Learning
- Neural Estimators for Conditional Mutual Information Using Nearest Neighbors Sampling
- Diversity regularization in deep ensembles
- Deep Ensembles on a Fixed Memory Budget: One Wide Network or Several Thinner Ones?
- StackOverflow vs Kaggle: A Study of Developer Discussions About Data Science