Controlling the false discovery rate via knockoffs
arXiv:1404.5609 · doi:10.1214/15-AOS1337
Abstract
In many fields of science, we observe a response variable together with a large number of potential explanatory variables, and would like to be able to discover which variables are truly associated with the response. At the same time, we need to know that the false discovery rate (FDR) - the expected fraction of false discoveries among all discoveries - is not too high, in order to assure the scientist that most of the discoveries are indeed true and replicable. This paper introduces the knockoff filter, a new variable selection procedure controlling the FDR in the statistical linear model whenever there are at least as many observations as variables. This method achieves exact FDR control in finite sample settings no matter the design or covariates, the number of variables in the model, or the amplitudes of the unknown regression coefficients, and does not require any knowledge of the noise level. As the name suggests, the method operates by manufacturing knockoff variables that are cheap - their construction does not require any new data - and are designed to mimic the correlation structure found within the existing variables, in a way that allows for accurate FDR control, beyond what is possible with permutation-based methods. The method of knockoffs is very general and flexible, and can work with a broad class of test statistics. We test the method in combination with statistics from the Lasso for sparse regression, and obtain empirical results showing that the resulting method has far more power than existing selection rules when the proportion of null variables is high.
Published at http://dx.doi.org/10.1214/15-AOS1337 in the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (5)
- Nearly unbiased variable selection under minimax concave penalty
- The sparsity and bias of the Lasso selection in high-dimensional linear regression
- Near-ideal model selection by minimization
- Exact and asymptotically robust permutation tests
- Stability Approach to Regularization Selection (StARS) for High Dimensional Graphical Models
Cited by in corpus (49)
- Machine Learning for Fluid Mechanics
- SLOPE - Adaptive variable selection via convex optimization
- High-Dimensional Inference: Confidence Intervals, -Values and R-Software hdi
- Gene Hunting with Knockoffs for Hidden Markov Models
- Deep Knockoffs
- Model-free Feature Screening and FDR Control with Knockoff Features
- Covariate powered cross-weighted multiple testing
- Metropolized Knockoff Sampling
- Causal Inference in Genetic Trio Studies
- High-dimensional inference in misspecified linear models
- STAR: A general interactive framework for FDR control under structural constraints
- Simultaneous high-probability bounds on the false discovery proportion in structured, regression, and online settings
- Interpreting Black Box Models via Hypothesis Testing
- PPFS: Predictive Permutation Feature Selection
- Rejoinder: "Gene Hunting with Hidden Markov Model Knockoffs"
- False Discovery Rate Controlled Heterogeneous Treatment Effect Detection for Online Controlled Experiments
- An Assumption-Free Exact Test For Fixed-Design Linear Models With Exchangeable Errors
- Non-convex Global Minimization and False Discovery Rate Control for the TREX
- Controlling the False Discovery Rate in Transformational Sparsity: Split Knockoffs
- Variable Selection with ABC Bayesian Forests
- Interpretable Classification of Bacterial Raman Spectra with Knockoff Wavelets
- Controlling Costs: Feature Selection on a Budget
- A powerful procedure that controls the false discovery rate with directional information
- CoxKnockoff: Controlled Feature Selection for the Cox Model Using Knockoffs
- ADAGES: adaptive aggregation with stability for distributed feature selection
- Feature-specific inference for penalized regression using local false discovery rates
- Controlling the False Discovery Rate via Competition: is the +1 needed?
- A Prototype Knockoff Filter for Group Selection with FDR Control
- Solving FDR-Controlled Sparse Regression Problems with Five Million Variables on a Laptop
- Aggregating Knockoffs for False Discovery Rate Control with an Application to Gut Microbiome Data
- Two-Stage Robust and Sparse Distributed Statistical Inference for Large-Scale Data
- Statistical hypothesis testing versus machine-learning binary classification: distinctions and guidelines
- Large-Scale Multiple Testing of Composite Null Hypotheses Under Heteroskedasticity
- FDP control in multivariate linear models using the bootstrap
- Residual permutation test for regression coefficient testing
- Replica Analysis for Ensemble Techniques in Variable Selection
- Variable Selection with the Knockoffs: Composite Null Hypotheses
- Flexible variable selection in the presence of missing data
- Simultaneous directional inference
- False Discovery Rate Control for Confounder Selection Using Mirror Statistics
- Inference for Large Panel Data with Many Covariates
- Cox reduction and confidence sets of models: a theoretical elucidation
- A Computationally Efficient Approach to False Discovery Rate Control and Power Maximisation via Randomisation and Mirror Statistic
- Learning to Increase the Power of Conditional Randomization Tests
- Co-Developing Causal Graphs with Domain Experts Guided by Weighted FDR-Adjusted p-values
- High-dimensional Statistical Inference and Variable Selection Using Sufficient Dimension Association
- A Pseudo Knockoff Filter for Correlated Features
- Can linear algebra create perfect knockoffs?
- Split Knockoffs for Multiple Comparisons: Controlling the Directional False Discovery Rate