Empirical Bayes inference in sparse high-dimensional generalized linear models
arXiv:2303.07854 · doi:10.1214/24-EJS2274
Abstract
High-dimensional linear models have been widely studied, but the developments in high-dimensional generalized linear models, or GLMs, have been slower. In this paper, we propose an empirical or data-driven prior leading to an empirical Bayes posterior distribution which can be used for estimation of and inference on the coefficient vector in a high-dimensional GLM, as well as for variable selection. We prove that our proposed posterior concentrates around the true/sparse coefficient vector at the optimal rate, provide conditions under which the posterior can achieve variable selection consistency, and prove a Bernstein--von Mises theorem that implies asymptotically valid uncertainty quantification. Computation of the proposed empirical Bayes posterior is simple and efficient, and is shown to perform well in simulations compared to existing Bayesian and non-Bayesian methods in terms of estimation and variable selection.
References in corpus (18)
- Optimal predictive model selection
- Sparsity information and regularization in the horseshoe and other shrinkage priors
- Bayesian linear regression with sparse priors
- Needles and Straw in a Haystack: Posterior concentration for possibly sparse sequences
- Calibrating general posterior credible regions
- Contributed Discussion to Uncertainty Quantification for the Horseshoe by Stéphanie van der Pas, Botond Szabó and Aad van der Vaart
- Empirical Bayes posterior concentration in sparse high-dimensional linear models
- Asymptotically minimax empirical Bayes estimation of a sparse normal mean vector
- Sparse Estimation by Exponential Weighting
- Kullback-Leibler aggregation and misspecified generalized linear models
- Gibbs posterior concentration rates under sub-exponential type losses
- Data-driven priors and their posterior concentration rates
- Empirical priors and coverage of posterior credible sets in a sparse normal mean model
- Empirical priors for prediction in sparse high-dimensional linear regression
- Empirical priors and posterior concentration rates for a monotone density
- Prediction risk for the horseshoe regression
- Polynomial Time and Private Learning of Unbounded Gaussian Mixture Models
- Advances in Bayesian model selection consistency for high-dimensional generalized linear models