Dual Extrapolation for Sparse Generalized Linear Models
arXiv:1907.05830
Abstract
Generalized Linear Models (GLM) form a wide class of regression and classification models, where prediction is a function of a linear combination of the input variables. For statistical inference in high dimension, sparsity inducing regularizations have proven to be useful while offering statistical guarantees. However, solving the resulting optimization problems can be challenging: even for popular iterative algorithms such as coordinate descent, one needs to loop over a large number of variables. To mitigate this, techniques known as screening rules and working sets diminish the size of the optimization problem at hand, either by progressively removing variables, or by solving a growing sequence of smaller problems. For both techniques, significant variables are identified thanks to convex duality arguments. In this paper, we show that the dual iterates of a GLM exhibit a Vector AutoRegressive (VAR) behavior after sign identification, when the primal problem is solved with proximal gradient descent or cyclic coordinate descent. Exploiting this regularity, one can construct dual points that offer tighter certificates of optimality, enhancing the performance of screening rules and helping to design competitive working set algorithms.
References in corpus (7)
- Pathwise coordinate optimization
- Safe Feature Elimination in Sparse Supervised Learning
- Faster Coordinate Descent via Adaptive Importance Sampling
- Local Convergence Properties of SAGA/Prox-SVRG and Acceleration
- Efficient Greedy Coordinate Descent for Composite Problems
- From safe screening rules to working sets for faster Lasso-type solvers
- A Fast, Principled Working Set Algorithm for Exploiting Piecewise Linear Structure in Convex Problems
Cited by in corpus (5)
- Implicit differentiation of Lasso-type models for hyperparameter optimization
- Statistical control for spatio-temporal MEG/EEG source imaging with desparsified multi-task Lasso
- Screening Rules and its Complexity for Active Set Identification
- Model identification and local linear convergence of coordinate descent
- On Newton Screening