L1-Penalization for Mixture Regression Models
arXiv:1202.6046 · doi:10.1007/s11749-010-0197-z
Abstract
We consider a finite mixture of regressions (FMR) model for high-dimensional inhomogeneous data where the number of covariates may be much larger than sample size. We propose an l1-penalized maximum likelihood estimator in an appropriate parameterization. This kind of estimation belongs to a class of problems where optimization and theory for non-convex functions is needed. This distinguishes itself very clearly from high-dimensional estimation with convex loss- or objective functions, as for example with the Lasso in linear or generalized linear models. Mixture models represent a prime and important example where non-convexity arises. For FMR models, we develop an efficient EM algorithm for numerical optimization with provable convergence properties. Our penalized estimator is numerically better posed (e.g., boundedness of the criterion function) than unpenalized maximum likelihood estimation, and it allows for effective statistical regularization including variable selection. We also present some asymptotic theory and oracle inequalities: due to non-convexity of the negative log-likelihood function, different mathematical arguments are needed than for problems with convex losses. Finally, we apply the new method to both simulated and real data.
This is the author's version of the work (published as a discussion paper in TEST, 2010, Volume 19, 209--285). The final publication is available at http://www.springerlink.com
References in corpus (9)
- Nearly unbiased variable selection under minimax concave penalty
- Pathwise coordinate optimization
- The sparsity and bias of the Lasso selection in high-dimensional linear regression
- Lasso-type recovery of sparse representations for high-dimensional data
- High-dimensional generalized linear models and the lasso
- Near-ideal model selection by minimization
- Sparsity oracle inequalities for the Lasso
- The Dantzig selector and sparsity oracle inequalities
- Some sharp performance bounds for least squares regression with regularization
Cited by in corpus (25)
- Challenges of Big Data Analysis
- SLOPE - Adaptive variable selection via convex optimization
- Robust subspace clustering
- On the Prediction Performance of the Lasso
- Estimation for High-Dimensional Linear Mixed-Effects Models Using -Penalization
- Endogeneity in high dimensions
- GLMMLasso: An Algorithm for High-Dimensional Generalized Linear Mixed Models Using L1-Penalization
- Nonconcave penalized composite conditional likelihood estimation of sparse Ising models
- Multi-species distribution modeling using penalized mixture of regressions
- Variance prior forms for high-dimensional Bayesian variable selection
- Estimating the error variance in a high-dimensional linear model
- Penalized estimation in high-dimensional hidden Markov models with state-specific graphical models
- Efficient Smoothed Concomitant Lasso Estimation for High Dimensional Regression
- A hierarchical Bayesian perspective on majorization-minimization for non-convex sparse regression: application to M/EEG source imaging
- Non-Concave Penalization in Linear Mixed-Effects Models and Regularized Selection of Fixed Effects
- Quasi-Likelihood and/or Robust Estimation in High Dimensions
- A two-part finite mixture quantile regression model for semi-continuous longitudinal data
- Supervised clustering of high dimensional data using regularized mixture modeling
- On estimation of the diagonal elements of a sparse precision matrix
- Robust Finite Mixture Regression for Heterogeneous Targets
- Semiparametric estimation of a two-component mixture of linear regressions in which one component is known
- Regression-based heterogeneity analysis to identify overlapping subgroup structure in high-dimensional data
- Estimating Heterogeneous Causal Effects of High-Dimensional Treatments: Application to Conjoint Analysis
- l1-Norm Minimization with Regula Falsi Type Root Finding Methods
- Compound Poisson Point Processes, Concentration and Oracle Inequalities