Feature Selection with Annealing for Computer Vision and Big Data Learning
arXiv:1310.2880 · doi:10.1109/TPAMI.2016.2544315
Abstract
Many computer vision and medical imaging problems are faced with learning from large-scale datasets, with millions of observations and features. In this paper we propose a novel efficient learning scheme that tightens a sparsity constraint by gradually removing variables based on a criterion and a schedule. The attractive fact that the problem size keeps dropping throughout the iterations makes it particularly suitable for big data learning. Our approach applies generically to the optimization of any differentiable loss function, and finds applications in regression, classification and ranking. The resultant algorithms build variable screening into estimation and are extremely simple to implement. We provide theoretical guarantees of convergence and selection consistency. In addition, one dimensional piecewise linear response functions are used to account for nonlinearity and a second order prior is imposed on these functions to avoid overfitting. Experiments on real and synthetic data show that the proposed method compares very well with other state of the art methods in regression, classification and ranking while being computationally very efficient and scalable.
18 pages, 9 figures
References in corpus (8)
- Nearly unbiased variable selection under minimax concave penalty
- Coordinate descent algorithms for nonconvex penalized regression, with applications to biological feature selection
- Variable selection in nonparametric additive models
- Piecewise linear regularized solution paths
- Sparsity oracle inequalities for the Lasso
- Sparse Online Learning via Truncated Gradient
- Group Iterative Spectrum Thresholding for Super-Resolution Sparse Spectral Selection
- Learning Topology and Dynamics of Large Recurrent Neural Networks
Cited by in corpus (8)
- Feature Selection with Annealing for Forecasting Financial Time Series
- Estimating Feature-Label Dependence Using Gini Distance Statistics
- Are screening methods useful in feature selection? An empirical study
- Simple Stochastic Gradient Methods for Non-Smooth Non-Convex Regularized Optimization
- A Novel Framework for Online Supervised Learning with Feature Selection
- Indirect Gaussian Graph Learning beyond Gaussianity
- Breaking the Limits in Urban Video Monitoring: Massive Crowd Sourced Surveillance over Vehicles
- A study of local optima for learning feature interactions using neural networks