SLOPE - Adaptive variable selection via convex optimization
arXiv:1407.3824 · doi:10.1214/15-AOAS842
Abstract
We introduce a new estimator for the vector of coefficients in the linear model , where has dimensions with possibly larger than . SLOPE, short for Sorted L-One Penalized Estimation, is the solution to \[\min_{b\in\mathbb{R}^p}\frac{1}{2}\Vert y-Xb\Vert _{\ell_2}^2+λ_1\vert b\vert _{(1)}+λ_2\vert b\vert_{(2)}+\cdots+λ_p\vert b\vert_{(p)},\] where and are the decreasing absolute values of the entries of . This is a convex program and we demonstrate a solution algorithm whose computational complexity is roughly comparable to that of classical procedures such as the Lasso. Here, the regularizer is a sorted norm, which penalizes the regression coefficients according to their rank: the higher the rank - that is, stronger the signal - the larger the penalty. This is similar to the Benjamini and Hochberg [J. Roy. Statist. Soc. Ser. B 57 (1995) 289-300] procedure (BH) which compares more significant -values with more stringent thresholds. One notable choice of the sequence is given by the BH critical values , where and is the quantile of a standard normal distribution. SLOPE aims to provide finite sample guarantees on the selected model; of special interest is the false discovery rate (FDR), defined as the expected proportion of irrelevant regressors among all selected predictors. Under orthogonal designs, SLOPE with provably controls FDR at level . Moreover, it also appears to have appreciable inferential properties under more general designs while having substantial power, as demonstrated in a series of experiments running on both simulated and real data.
Published at http://dx.doi.org/10.1214/15-AOAS842 in the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (8)
- Controlling the false discovery rate via knockoffs
- High-dimensional variable selection
- SLOPE - Adaptive variable selection via convex optimization
- L1-Penalization for Mixture Regression Models
- A simple forward selection procedure based on false discovery rate control
- P-values for high-dimensional regression
- False Variable Selection Rates in Regression
- A model selection approach to genome wide association studies
Cited by in corpus (75)
- A Differential Equation for Modeling Nesterov's Accelerated Gradient Method: Theory and Insights
- SLOPE - Adaptive variable selection via convex optimization
- High-Dimensional Inference: Confidence Intervals, -Values and R-Software hdi
- Sparse Portfolio Selection via the sorted -Norm
- Spike-and-Slab Meets LASSO: A Review of the Spike-and-Slab LASSO
- Panning for Gold: Model-X Knockoffs for High-dimensional Controlled Variable Selection
- Nonconvex penalties with analytical solutions for one-bit compressive sensing
- On Online Control of False Discovery Rate
- Sparse Estimation with Strongly Correlated Variables using Ordered Weighted L1 Regularization
- Predicting the redshift of gamma-ray loud AGNs using Supervised Machine Learning: Part 2
- Regularization and the small-ball method II: complexity dependent error rates
- Asymptotic Confidence Regions for High-dimensional Structured Sparsity
- The Trimmed Lasso: Sparsity and Robustness
- SLOPE is Adaptive to Unknown Sparsity and Asymptotically Minimax
- Differentially Private False Discovery Rate Control
- False Discovery Rate Control via Data Splitting
- Estimation bounds and sharp oracle inequalities of regularized procedures with Lipschitz loss functions
- Fast OSCAR and OWL Regression via Safe Screening Rules
- Learning from MOM's principles: Le Cam's approach
- Asymptotic Analysis via Stochastic Differential Equations of Gradient Descent Algorithms in Statistical and Computational Paradigms
- A Scalable Empirical Bayes Approach to Variable Selection in Generalized Linear Models
- Sharp Oracle Inequalities for Square Root Regularization
- Optimal link prediction with matrix logistic regression
- Interplay of minimax estimation and minimax support recovery under sparsity
- The Geometry of Uniqueness, Sparsity and Clustering in Penalized Estimation
- Parameter Estimation with the Ordered Regularization via an Alternating Direction Method of Multipliers
- Feasibility and A Fast Algorithm for Euclidean Distance Matrix Optimization with Ordinal Constraints
- Adaptive Bayesian SLOPE -- High-dimensional Model Selection with Missing Values
- High-dimensional semi-supervised learning: in search for optimal inference of the mean
- Fast Signal Recovery from Saturated Measurements by Linear Loss and Nonconvex Penalties
- Fast Saddle-Point Algorithm for Generalized Dantzig Selector and FDR Control with the Ordered l1-Norm
- The FDR-Linking Theorem
- High dimensional regression and matrix estimation without tuning parameters
- Towards Optimal Problem Dependent Generalization Error Bounds in Statistical Learning Theory
- Solving L1-regularized SVMs and related linear programs: Revisiting the effectiveness of Column and Constraint Generation
- Optimal False Discovery Control of Minimax Estimator
- The Strong Screening Rule for SLOPE
- On the Exponentially Weighted Aggregate with the Laplace Prior
- Genetic variant selection: learning across traits and sites
- Structure Learning of Gaussian Markov Random Fields with False Discovery Rate Control
- Empirical Bayes cumulative -value multiple testing procedure for sparse sequences
- Efficient Designs of SLOPE Penalty Sequences in Finite Dimension
- A Scalable Empirical Bayes Approach to Variable Selection
- Efficient Projection Algorithms onto the Weighted l1 Ball
- Familywise Error Rate Control via Knockoffs
- Error bounds for sparse classifiers in high-dimensions
- Outlier-robust sparse/low-rank least-squares regression and robust matrix completion
- Variable Selection with Second-Generation P-Values
- M-estimation with the Trimmed l1 Penalty
- Variable Selection via Adaptive False Negative Control in Linear Regression
- On Regularized Square-root Regression Problems: Distributionally Robust Interpretation and Fast Computations
- Minimax Optimal Estimation in Partially Linear Additive Models under High Dimension
- Efficient Path Algorithms for Clustered Lasso and OSCAR
- Sparse Data-Driven Random Projection in Regression for High-Dimensional Data
- Multiple Regression for Matrix and Vector Predictors: Models, Theory, Algorithms, and Beyond
- Double spike Dirichlet priors for structured weighting
- Improved bounds for Square-Root Lasso and Square-Root Slope
- The Complete Lasso Tradeoff Diagram
- A Unified Framework for Pattern Recovery in Penalized and Thresholded Estimation and its Geometry
- Fast projection onto the ordered weighted norm ball
- Improved error rates for sparse (group) learning with Lipschitz loss functions
- Iteratively Reweighted -Penalized Robust Regression
- Deep-gKnock: nonlinear group-feature selection with deep neural network
- Constrained Machine Learning: The Bagel Framework
- Selection and Estimation Optimality in High Dimensions with the TWIN Penalty
- Efficient Predictor Ranking and False Discovery Proportion Control in High-Dimensional Regression
- A note on sharp oracle bounds for Slope and Lasso
- An Exact Solution Path Algorithm for SLOPE and Quasi-Spherical OSCAR
- DebiNet: Debiasing Linear Models with Nonlinear Overparameterized Neural Networks
- Robust selection of predictors and conditional outlier detection in a perturbed large-dimensional regression context
- Characterizing the SLOPE Trade-off: A Variational Perspective and the Donoho-Tanner Limit
- Finding Statistically Significant Interactions between Continuous Features
- Sharp Oracle Inequalities for Low-complexity Priors
- Nested Model Averaging on Solution Path for High-dimensional Linear Regression
- Learning Games and Rademacher Observations Losses