Weight Uncertainty in Neural Networks
arXiv:1505.05424
Abstract
We introduce a new, efficient, principled and backpropagation-compatible algorithm for learning a probability distribution on the weights of a neural network, called Bayes by Backprop. It regularises the weights by minimising a compression cost, known as the variational free energy or the expected lower bound on the marginal likelihood. We show that this principled kind of regularisation yields comparable performance to dropout on MNIST classification. We then demonstrate how the learnt uncertainty in the weights can be used to improve generalisation in non-linear regression problems, and how this weight uncertainty can be used to drive the exploration-exploitation trade-off in reinforcement learning.
In Proceedings of the 32nd International Conference on Machine Learning (ICML 2015)
References in corpus (3)
Cited by in corpus (49)
- Deep and Confident Prediction for Time Series at Uber
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- Peer-to-peer Federated Learning on Graphs
- Parallel and Distributed Thompson Sampling for Large-scale Accelerated Exploration of Chemical Space
- Quality of Uncertainty Quantification for Bayesian Neural Network Inference
- An Equivalence of Fully Connected Layer and Convolutional Layer
- Model Selection in Bayesian Neural Networks via Horseshoe Priors
- 'In-Between' Uncertainty in Bayesian Neural Networks
- Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation
- Uncertainty Decomposition in Bayesian Neural Networks with Latent Variables
- Subspace Inference for Bayesian Deep Learning
- Deep Robust Kalman Filter
- Bayesian Semisupervised Learning with Deep Generative Models
- Uncertainty quantification of molecular property prediction with Bayesian neural networks
- Bayesian Policy Gradients via Alpha Divergence Dropout Inference
- Efficient exploration with Double Uncertain Value Networks
- Expressive Priors in Bayesian Neural Networks: Kernel Combinations and Periodic Functions
- Variational Inference to Measure Model Uncertainty in Deep Neural Networks
- Bayesian Learning of Neural Network Architectures
- Meta-Learning surrogate models for sequential decision making
- Bayesian Neural Networks
- Tied Hidden Factors in Neural Networks for End-to-End Speaker Recognition
- Radial and Directional Posteriors for Bayesian Neural Networks
- A Convergence Analysis for A Class of Practical Variance-Reduction Stochastic Gradient MCMC
- Stochastic Maximum Likelihood Optimization via Hypernetworks
- A Scale Mixture Perspective of Multiplicative Noise in Neural Networks
- Continual Learning in Deep Neural Network by Using a Kalman Optimiser
- On architectural choices in deep learning: From network structure to gradient convergence and parameter estimation
- Fast learning rate of deep learning via a kernel perspective
- Comparing Semi-Parametric Model Learning Algorithms for Dynamic Model Estimation in Robotics
- Ensemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network
- Trust Region Value Optimization using Kalman Filtering
- Energy Confused Adversarial Metric Learning for Zero-Shot Image Retrieval and Clustering
- Applying SVGD to Bayesian Neural Networks for Cyclical Time-Series Prediction and Inference
- Adaptive Temporal-Difference Learning for Policy Evaluation with Per-State Uncertainty Estimates
- A Generative Model for Sampling High-Performance and Diverse Weights for Neural Networks
- Bayesian Neural Networks at Finite Temperature
- Combining Model and Parameter Uncertainty in Bayesian Neural Networks
- Adaptively Preconditioned Stochastic Gradient Langevin Dynamics
- Radial Prediction Layer
- Variational Bayes: A report on approaches and applications
- Neural Likelihoods for Multi-Output Gaussian Processes
- Probabilistic Discriminative Learning with Layered Graphical Models
- Uncertainty quantification of molecular property prediction using Bayesian neural network models
- Stochastic Gradient MCMC with Stale Gradients
- Revisiting hard thresholding for DNN pruning
- Performance Measurement for Deep Bayesian Neural Network
- Nearest-Neighbor Neural Networks for Geostatistics
- Robustness Against Outliers For Deep Neural Networks By Gradient Conjugate Priors