Sparsity information and regularization in the horseshoe and other shrinkage priors
arXiv:1707.01694 · doi:10.1214/17-EJS1337SI
Abstract
The horseshoe prior has proven to be a noteworthy alternative for sparse Bayesian estimation, but has previously suffered from two problems. First, there has been no systematic way of specifying a prior for the global shrinkage hyperparameter based on the prior information about the degree of sparsity in the parameter vector. Second, the horseshoe prior has the undesired property that there is no possibility of specifying separately information about sparsity and the amount of regularization for the largest coefficients, which can be problematic with weakly identified parameters, such as the logistic regression coefficients in the case of data separation. This paper proposes solutions to both of these problems. We introduce a concept of effective number of nonzero parameters, show an intuitive way of formulating the prior for the global hyperparameter based on the sparsity assumptions, and argue that the previous default choices are dubious based on their tendency to favor solutions with more unshrunk parameters than we typically expect a priori. Moreover, we introduce a generalization to the horseshoe prior, called the regularized horseshoe, that allows us to specify a minimum level of regularization to the largest values. We show that the new prior can be considered as the continuous counterpart of the spike-and-slab prior with a finite slab width, whereas the original horseshoe resembles the spike-and-slab with an infinitely wide slab. Numerical experiments on synthetic and real world data illustrate the benefit of both of these theoretical advances.
References in corpus (2)
Cited by in corpus (42)
- Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors
- Leave-One-Out Cross-Validation for Bayesian Model Comparison in Large Data
- Practical Hilbert space approximate Bayesian Gaussian processes for probabilistic programming
- Intuitive Joint Priors for Bayesian Linear Multilevel Models: The R2D2M2 prior
- Projection Predictive Inference for Generalized Linear and Additive Multilevel Models
- Some models are useful, but how do we know which ones? Towards a unified Bayesian model taxonomy
- Pathfinder: Parallel quasi-Newton variational inference
- A fully Bayesian sparse polynomial chaos expansion approach with joint priors on the coefficients and global selection of terms
- Horseshoe Prior Bayesian Quantile Regression
- Enforcing stationarity through the prior in vector autoregressions
- The Reciprocal Bayesian LASSO
- GOES GLM, Biased Bolides, and Debiased Distributions
- Quantifying sources of uncertainty in drug discovery predictions with probabilistic models
- Hamiltonian Monte Carlo using an adjoint-differentiated Laplace approximation: Bayesian inference for latent Gaussian models and beyond
- Efficient estimation and correction of selection-induced bias with order statistics
- Adaptive Uncertainty-Guided Model Selection for Data-Driven PDE Discovery
- Inconsistency identification in network meta-analysis via stochastic search variable selection
- Latent space projection predictive inference
- Bayesian sparsification for deep neural networks with Bayesian model reduction
- Proximal MCMC for Bayesian Inference of Constrained and Regularized Estimation
- Empirical Bayes inference in sparse high-dimensional generalized linear models
- Translating predictive distributions into informative priors
- Projection predictive variable selection for discrete response families with finite support
- : A robust MCMC convergence diagnostic with uncertainty using decision tree classifiers
- Informative Bayesian Neural Network Priors for Weak Signals
- Probabilistically-autoencoded horseshoe-disentangled multidomain item-response theory models
- Adaptive Path Sampling in Metastable Posterior Distributions
- A Laplace Mixture Representation of the Horseshoe and Some Implications
- Bayesian Dynamical System Identification With Unified Sparsity Priors And Model Uncertainty
- Generalized Decomposition Priors on R2
- Sparse encoding for more-interpretable feature-selecting representations in probabilistic matrix factorization
- Design and Structure Dependent Priors for Scale Parameters in Latent Gaussian Models
- Empirical Bayesian Inference using Joint Sparsity
- Mapping poverty at multiple geographical scales
- Time Fused Coefficient SIR Model with Application to COVID-19 Epidemic in the United States
- A fast asynchronous MCMC sampler for sparse Bayesian inference
- Bayesian Regularization: From Tikhonov to Horseshoe
- HALO: Learning to Prune Neural Networks with Shrinkage
- The ARR2 prior: flexible predictive prior definition for Bayesian auto-regressions
- Encoding Domain Information with Sparse Priors for Inferring Explainable Latent Variables
- Gibbs Sampling using Anti-correlation Gaussian Data Augmentation, with Applications to L1-ball-type Models
- Horseshoe Forests for High-Dimensional Causal Survival Analysis