The influence of feature selection methods on accuracy, stability and interpretability of molecular signatures
arXiv:1101.5008 · doi:10.1371/journal.pone.0028210
Abstract
Motivation: Biomarker discovery from high-dimensional data is a crucial problem with enormous applications in biology and medicine. It is also extremely challenging from a statistical viewpoint, but surprisingly few studies have investigated the relative strengths and weaknesses of the plethora of existing feature selection methods. Methods: We compare 32 feature selection methods on 4 public gene expression datasets for breast cancer prognosis, in terms of predictive performance, stability and functional interpretability of the signatures they produce. Results: We observe that the feature selection method has a significant influence on the accuracy, stability and interpretability of signatures. Simple filter methods generally outperform more complex embedded or wrapper methods, and ensemble feature selection has generally no positive effect. Overall a simple Student's t-test seems to provide the best results. Availability: Code and data are publicly available at http://cbio.ensmp.fr/~ahaury/.
Cited by in corpus (17)
- Correlation and variable importance in random forests
- Robust and Complex Approach of Pathological Speech Signal Analysis
- Gains in Power from Structured Two-Sample Tests of Means on Graphs
- More power via graph-structured tests for differential expression of gene networks
- "Dave...I can assure you...that it's going to be all right..." -- A definition, case for, and survey of algorithmic assurances in human-autonomy trust relationships
- Small-sample Brain Mapping: Sparse Recovery on Spatially Correlated Designs with Randomization and Clustering
- Methods and Models for Interpretable Linear Classification
- High-Dimensional Feature Selection for Genomic Datasets
- A Novel Approach for Stable Selection of Informative Redundant Features from High Dimensional fMRI Data
- Distinguishing Cell Phenotype Using Cell Epigenotype
- Network-Guided Biomarker Discovery
- Select and Attend: Towards Controllable Content Selection in Text Generation
- Fast acquisition of spin-wave dispersion by compressed sensing
- A Pipeline for Integrated Theory and Data-Driven Modeling of Genomic and Clinical Data
- Filter Methods for Feature Selection in Supervised Machine Learning Applications -- Review and Benchmark
- Testing for Feature Relevance: The HARVEST Algorithm
- iGPSe: A Visual Analytic System for Integrative Genomic Based Cancer Patient Stratification