Feature Selection via Regularized Trees
arXiv:1201.1587
Abstract
We propose a tree regularization framework, which enables many tree models to perform feature selection efficiently. The key idea of the regularization framework is to penalize selecting a new feature for splitting when its gain (e.g. information gain) is similar to the features used in previous splits. The regularization framework is applied on random forest and boosted trees here, and can be easily applied to other tree models. Experimental studies show that the regularized trees can select high-quality feature subsets with regard to both strong and weak classifiers. Because tree models can naturally deal with categorical and numerical variables, missing values, different scales between variables, interactions and nonlinearities etc., the tree regularization framework provides an effective and efficient feature selection solution for many practical problems.
8 pages; The 2012 International Joint Conference on Neural Networks (IJCNN), IEEE, 2012
References in corpus (1)
Cited by in corpus (8)
- Big Data Analytics for Dynamic Energy Management in Smart Grids
- Variable selection for BART: An application to gene regulation
- Feature selection revisited in the single-cell era
- Approximate False Positive Rate Control in Selection Frequency for Random Forest
- A new parsimonious method for classifying Cancer Tissue-of-Origin Based on DNA Methylation 450K data
- Gene selection with guided regularized random forest
- Non-subjective power analysis to detect G*E interactions in Genome-Wide Association Studies in presence of confounding factor
- Robustness of Random Forest-based gene selection methods