Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks
arXiv:1511.06314
Abstract
Convolutional Neural Networks have achieved state-of-the-art performance on a wide range of tasks. Most benchmarks are led by ensembles of these powerful learners, but ensembling is typically treated as a post-hoc procedure implemented by averaging independently trained models with model variation induced by bagging or random initialization. In this paper, we rigorously treat ensembling as a first-class problem to explicitly address the question: what are the best strategies to create an ensemble? We first compare a large number of ensembling strategies, and then propose and evaluate novel strategies, such as parameter sharing (through a new family of models we call TreeNets) as well as training under ensemble-aware and diversity-encouraging losses. We demonstrate that TreeNets can improve ensemble performance and that diverse ensembles can be trained end-to-end under a unified loss, achieving significantly higher "oracle" accuracies than classical ensembles.
Cited by in corpus (66)
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- Deep Ensembles: A Loss Landscape Perspective
- Diversity in Machine Learning
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Speed/accuracy trade-offs for modern convolutional object detectors
- A Probabilistic U-Net for Segmentation of Ambiguous Images
- High-Quality Prediction Intervals for Deep Learning: A Distribution-Free, Ensembled Approach
- Deep Learning for Brain Age Estimation: A Systematic Review
- Stochastic Segmentation Networks: Modelling Spatially Correlated Aleatoric Uncertainty
- Artificial Neural Networks Modelling of Wall Pressure Spectra Beneath Turbulent Boundary Layers
- Convolutional neural networks that teach microscopes how to image
- Quality Resilient Deep Neural Networks
- Optimized ensemble deep learning framework for scalable forecasting of dynamics containing extreme events
- Hyperparameter Ensembles for Robustness and Uncertainty Quantification
- Uncertainty Quantification and Deep Ensembles
- Training independent subnetworks for robust prediction
- A Hierarchical Probabilistic U-Net for Modeling Multi-Scale Ambiguities
- Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition
- Out-of-Distribution Detection using Multiple Semantic Label Representations
- Data-Driven Selection and Spectral Classification of White Dwarf Stars
- Attention-based Ensemble for Deep Metric Learning
- Coarse to Fine Multi-Resolution Temporal Convolutional Network
- TMS-Net: A Segmentation Network Coupled With A Run-time Quality Control Method For Robust Cardiac Image Segmentation
- Bayesian Inference with Anchored Ensembles of Neural Networks, and Application to Exploration in Reinforcement Learning
- Neural Ensemble Search for Uncertainty Estimation and Dataset Shift
- On Last-Layer Algorithms for Classification: Decoupling Representation from Uncertainty Estimation
- Objects are Different: Flexible Monocular 3D Object Detection
- EnsembleNet: End-to-End Optimization of Multi-headed Models
- Group Ensemble: Learning an Ensemble of ConvNets in a single ConvNet
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Deep Ensembles for Low-Data Transfer Learning
- Cultivating DNN Diversity for Large Scale Video Labelling
- Should attention be all we need? The epistemic and ethical implications of unification in machine learning
- Semi-Supervised Deep Ensembles for Blind Image Quality Assessment
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System
- Evaluating Scalable Uncertainty Estimation Methods for DNN-Based Molecular Property Prediction
- Robust Semantic Segmentation with Superpixel-Mix
- Probabilistic Regression of Rotations using Quaternion Averaging and a Deep Multi-Headed Network
- Deep Ensembles on a Fixed Memory Budget: One Wide Network or Several Thinner Ones?
- Towards Oracle Knowledge Distillation with Neural Architecture Search
- Deep Anti-Regularized Ensembles provide reliable out-of-distribution uncertainty quantification
- Improving Confidence Estimates for Unfamiliar Examples
- TRADI: Tracking deep neural network weight distributions for uncertainty estimation
- Aggressive Q-Learning with Ensembles: Achieving Both High Sample Efficiency and High Asymptotic Performance
- Why have a Unified Predictive Uncertainty? Disentangling it using Deep Split Ensembles
- Prediction intervals for Deep Neural Networks
- Accelerating Multi-Model Inference by Merging DNNs of Different Weights
- Deep Architectures and Ensembles for Semantic Video Classification
- Sparse MoEs meet Efficient Ensembles
- Uncertainty quantification for White Matter Hyperintensity segmentation detects silent failures and improves automated Fazekas quantification
- Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning
- Quantifying Uncertainty in Deep Spatiotemporal Forecasting
- RBUE: A ReLU-Based Uncertainty Estimation Method of Deep Neural Networks
- To Boost or not to Boost: On the Limits of Boosted Neural Networks
- Deep Extreme Feature Extraction: New MVA Method for Searching Particles in High Energy Physics
- IEA: Inner Ensemble Average within a convolutional neural network
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- A Variational View on Bootstrap Ensembles as Bayesian Inference
- Quantifying Epistemic Uncertainty in Deep Learning
- DiverseNet: When One Right Answer is not Enough
- JUWELS Booster -- A Supercomputer for Large-Scale AI Research
- Multi-headed Neural Ensemble Search
- Regularizing Neural Networks via Stochastic Branch Layers
- Structured Ensembles: an Approach to Reduce the Memory Footprint of Ensemble Methods
- Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression