Hierarchical Sparse Modeling: A Choice of Two Group Lasso Formulations
arXiv:1512.01631 · doi:10.1214/17-STS622
Abstract
Demanding sparsity in estimated models has become a routine practice in statistics. In many situations, we wish to require that the sparsity patterns attained honor certain problem-specific constraints. Hierarchical sparse modeling (HSM) refers to situations in which these constraints specify that one set of parameters be set to zero whenever another is set to zero. In recent years, numerous papers have developed convex regularizers for this form of sparsity structure, which arises in many areas of statistics including interaction modeling, time series analysis, and covariance estimation. In this paper, we observe that these methods fall into two frameworks, the group lasso (GL) and latent overlapping group lasso (LOG), which have not been systematically compared in the context of HSM. The purpose of this paper is to provide a side-by-side comparison of these two frameworks for HSM in terms of their statistical properties and computational efficiency. We call special attention to GL's more aggressive shrinkage of parameters deep in the hierarchy, a property not shared by LOG. In terms of computation, we introduce a finite-step algorithm that exactly solves the proximal operator of LOG for a certain simple HSM structure; we later exploit this to develop a novel path-based block coordinate descent scheme for general HSM structures. Both algorithms greatly improve the computational performance of LOG. Finally, we compare the two methods in the context of covariance estimation, where we introduce a new sparsely-banded estimator using LOG, which we show achieves the statistical advantages of an existing GL-based method but is simpler to express and more efficient to compute.
30 pages, 13 figures
References in corpus (6)
- The composite absolute penalties family for grouped and hierarchical variable selection
- Exploring Large Feature Spaces with Hierarchical Multiple Kernel Learning
- Sparse estimation of large covariance matrices via a nested Lasso penalty
- Structured variable selection and estimation
- Generalized Additive Model Selection
- Convex Modeling of Interactions with Strong Heredity
Cited by in corpus (8)
- Learning Hierarchical Interactions at Scale: A Convex Optimization Approach
- Estimating Heterogeneous Causal Effects of High-Dimensional Treatments: Application to Conjoint Analysis
- A first-order optimization algorithm for statistical learning with hierarchical sparsity structure
- Testing for Conditional Mean Independence with Covariates through Martingale Difference Divergence
- Fitting ARMA Time Series Models without Identification: A Proximal Approach
- Reluctant generalized additive modeling
- A likelihood-based approach for multivariate categorical response regression in high dimensions
- Predicting Census Survey Response Rates With Parsimonious Additive Models and Structured Interactions