Accelerated Gradient Methods for Sparse Statistical Learning with Nonconvex Penalties
arXiv:2009.10629 · doi:10.1007/s11222-023-10371-8
Abstract
Nesterov's accelerated gradient (AG) is a popular technique to optimize objective functions comprising two components: a convex loss and a penalty function. While AG methods perform well for convex penalties, such as the LASSO, convergence issues may arise when it is applied to nonconvex penalties, such as SCAD. A recent proposal generalizes Nesterov's AG method to the nonconvex setting. The proposed algorithm requires specification of several hyperparameters for its practical application. Aside from some general conditions, there is no explicit rule for selecting the hyperparameters, and how different selection can affect convergence of the algorithm. In this article, we propose a hyperparameter setting based on the complexity upper bound to accelerate convergence, and consider the application of this nonconvex AG algorithm to high-dimensional linear and logistic sparse learning problems. We further establish the rate of convergence and present a simple and useful bound to characterize our proposed optimal damping sequence. Simulation studies show that convergence can be made, on average, considerably faster than that of the conventional proximal gradient algorithm. Our experiments also show that the proposed method generally outperforms the current state-of-the-art methods in terms of signal recovery.
42 pages, 13 figures
References in corpus (6)
- Nearly unbiased variable selection under minimax concave penalty
- Pathwise coordinate optimization
- Coordinate descent algorithms for nonconvex penalized regression, with applications to biological feature selection
- Calibrating nonconvex penalized regression in ultra-high dimension
- Non-Concave Penalization in Linear Mixed-Effects Models and Regularized Selection of Fixed Effects
- Strong rules for nonconvex penalties and their implications for efficient algorithms in high-dimensional regression