Informative Bayesian Neural Network Priors for Weak Signals
arXiv:2002.10243 · doi:10.1214/21-BA1291
Abstract
Encoding domain knowledge into the prior over the high-dimensional weight space of a neural network is challenging but essential in applications with limited data and weak signals. Two types of domain knowledge are commonly available in scientific applications: 1. feature sparsity (fraction of features deemed relevant); 2. signal-to-noise ratio, quantified, for instance, as the proportion of variance explained (PVE). We show how to encode both types of domain knowledge into the widely used Gaussian scale mixture priors with Automatic Relevance Determination. Specifically, we propose a new joint prior over the local (i.e., feature-specific) scale parameters that encodes knowledge about feature sparsity, and a Stein gradient optimization to tune the hyperparameters in such a way that the distribution induced on the model's PVE matches the prior distribution. We show empirically that the new prior improves prediction accuracy, compared to existing neural network priors, on several publicly available datasets and in a genetics application where signals are weak and sparse, often outperforming even computationally intensive cross-validation for hyperparameter tuning.
25 pages, 8 figures, 4 tables
References in corpus (12)
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Sparsity information and regularization in the horseshoe and other shrinkage priors
- A General Framework for the Parametrization of Hierarchical Models
- Deep Learning: A Bayesian Perspective
- The Horseshoe Estimator: Posterior Concentration around Nearly Black Vectors
- What Are Bayesian Neural Network Posteriors Really Like?
- Model Selection in Bayesian Neural Networks via Horseshoe Priors
- Reversible Jump MCMC Simulated Annealing for Neural Networks
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks
- Expressive Priors in Bayesian Neural Networks: Kernel Combinations and Periodic Functions