Innovated interaction screening for high-dimensional nonlinear classification
arXiv:1501.01029 · doi:10.1214/14-AOS1308
Abstract
This paper is concerned with the problems of interaction screening and nonlinear classification in a high-dimensional setting. We propose a two-step procedure, IIS-SQDA, where in the first step an innovated interaction screening (IIS) approach based on transforming the original -dimensional feature vector is proposed, and in the second step a sparse quadratic discriminant analysis (SQDA) is proposed for further selecting important interactions and main effects and simultaneously conducting classification. Our IIS approach screens important interactions by examining only features instead of all two-way interactions of order . Our theory shows that the proposed method enjoys sure screening property in interaction selection in the high-dimensional setting of growing exponentially with the sample size. In the selection and classification step, we establish a sparse inequality on the estimated coefficient vector for QDA and prove that the classification error of our procedure can be upper-bounded by the oracle classification error plus some smaller order term. Extensive simulation studies and real data analysis show that our proposal compares favorably with existing methods in interaction selection and high-dimensional classification.
Published at http://dx.doi.org/10.1214/14-AOS1308 in the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (9)
- Nearly unbiased variable selection under minimax concave penalty
- Regularized estimation of large covariance matrices
- Sparse permutation invariant covariance estimation
- High-dimensional classification using features annealed independence rules
- High-dimensional generalized linear models and the lasso
- A unified approach to model selection and sparse recovery using regularized least squares
- Sparse linear discriminant analysis by thresholding for high dimensional data
- Honest variable selection in linear and logistic regression models via and penalization
- Structures and Assumptions: Strategies to Harness Gene Gene and Gene Environment Interactions in GWAS