Bayesian Dark Knowledge
arXiv:1506.04416
Abstract
We consider the problem of Bayesian parameter estimation for deep neural networks, which is important in problem settings where we may have little data, and/ or where we need accurate posterior predictive densities, e.g., for applications involving bandits or active learning. One simple approach to this is to use online Monte Carlo methods, such as SGLD (stochastic gradient Langevin dynamics). Unfortunately, such a method needs to store many copies of the parameters (which wastes memory), and needs to make predictions using many versions of the model (which wastes time). We describe a method for "distilling" a Monte Carlo approximation to the posterior predictive density into a more compact form, namely a single deep neural network. We compare to two very recent approaches to Bayesian neural networks, namely an approach based on expectation propagation [Hernandez-Lobato and Adams, 2015] and an approach based on variational Bayes [Blundell et al., 2015]. Our method performs better than both of these, is much simpler to implement, and uses less computation at test time.
final version submitted to NIPS 2015
References in corpus (5)
- Distilling the Knowledge in a Neural Network
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Weight Uncertainty in Neural Networks
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Bayesian Posterior Sampling via Stochastic Gradient Fisher Scoring
Cited by in corpus (73)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Hands-on Bayesian Neural Networks -- a Tutorial for Deep Learning Users
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- Self-training with Noisy Student improves ImageNet classification
- Multiplicative Normalizing Flows for Variational Bayesian Neural Networks
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Bayesian Model-Agnostic Meta-Learning
- Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
- Distilling Knowledge from Deep Networks with Applications to Healthcare Domain
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- Self-Distillation Amplifies Regularization in Hilbert Space
- Dynamic Sampling Networks for Efficient Action Recognition in Videos
- Task Agnostic Continual Learning Using Online Variational Bayes
- Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
- Self-Distillation as Instance-Specific Label Smoothing
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Hydra: Preserving Ensemble Diversity for Model Distillation
- Conditional Generative Moment-Matching Networks
- Adversarial Distillation of Bayesian Neural Network Posteriors
- Laplace Redux -- Effortless Bayesian Deep Learning
- Channel Compression: Rethinking Information Redundancy among Channels in CNN Architecture
- Approximate Inference with Amortised MCMC
- Advances in Variational Inference
- Inhibited Softmax for Uncertainty Estimation in Neural Networks
- Resource-Efficient Neural Networks for Embedded Systems
- Attended Temperature Scaling: A Practical Approach for Calibrating Deep Neural Networks
- A Survey on Bayesian Deep Learning
- Bayesian Inference for Large Scale Image Classification
- Task Agnostic Continual Learning Using Online Variational Bayes with Fixed-Point Updates
- Infinitely Deep Bayesian Neural Networks with Stochastic Differential Equations
- Diving deeper into mentee networks
- The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding
- Bayesian Graph Convolutional Neural Networks using Node Copying
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- OpenEI: An Open Framework for Edge Intelligence
- Shift-based Primitives for Efficient Convolutional Neural Networks
- A Variational Dirichlet Framework for Out-of-Distribution Detection
- Learning Deep Representations with Probabilistic Knowledge Transfer
- Accelerating Monte Carlo Bayesian Inference via Approximating Predictive Uncertainty over Simplex
- Bayesian Graph Convolutional Neural Networks Using Non-Parametric Graph Learning
- Optimizing for Interpretability in Deep Neural Networks with Tree Regularization
- Learning and Refining of Privileged Information-based RNNs for Action Recognition from Depth Sequences
- The Bayesian Learning Rule
- LiBRe: A Practical Bayesian Approach to Adversarial Detection
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Scalable Training of Inference Networks for Gaussian-Process Models
- URSABench: Comprehensive Benchmarking of Approximate Bayesian Inference Methods for Deep Neural Networks
- Generalized Bayesian Posterior Expectation Distillation for Deep Neural Networks
- Efficient Evaluation-Time Uncertainty Estimation by Improved Distillation
- AntMan: Sparse Low-Rank Compression to Accelerate RNN inference
- A First Look at Deep Learning Apps on Smartphones
- Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and Memorization
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- High-Order Stochastic Gradient Thermostats for Bayesian Learning of Deep Models
- On the Robustness of Active Learning
- Dirichlet Pruning for Neural Network Compression
- Efficient and Robust Machine Learning for Real-World Systems
- Accurate and Reliable Forecasting using Stochastic Differential Equations
- Assessing the Robustness of Bayesian Dark Knowledge to Posterior Uncertainty
- Scalable Approximate Inference and Some Applications
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
- The Bayesian Method of Tensor Networks
- Sampling Prediction-Matching Examples in Neural Networks: A Probabilistic Programming Approach
- Full-Stack Filters to Build Minimum Viable CNNs
- An Improving Framework of regularization for Network Compression
- Bayes-Adaptive Deep Model-Based Policy Optimisation
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- Risk factor identification for incident heart failure using neural network distillation and variable selection
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- Generative Particle Variational Inference via Estimation of Functional Gradients