Least Squares Revisited: Scalable Approaches for Multi-class Prediction
arXiv:1310.1949
Abstract
This work provides simple algorithms for multi-class (and multi-label) prediction in settings where both the number of examples n and the data dimension d are relatively large. These robust and parameter free algorithms are essentially iterative least-squares updates and very versatile both in theory and in practice. On the theoretical front, we present several variants with convergence guarantees. Owing to their effective use of second-order structure, these algorithms are substantially better than first-order methods in many practical scenarios. On the empirical side, we present a scalable stagewise variant of our approach, which achieves dramatic computational speedups over popular optimization packages such as Liblinear and Vowpal Wabbit on standard datasets (MNIST and CIFAR-10), while attaining state-of-the-art accuracies.
References in corpus (3)
Cited by in corpus (9)
- Scalable Kernel Methods via Doubly Stochastic Gradients
- Logarithmic Time Online Multiclass prediction
- Matrix Completion Under Monotonic Single Index Models
- Large Scale Kernel Learning using Block Coordinate Descent
- Simultaneous Learning of Trees and Representations for Extreme Classification and Density Estimation
- A Generalizable and Accessible Approach to Machine Learning with Global Satellite Imagery
- Provable Tensor Methods for Learning Mixtures of Generalized Linear Models
- Scalable Multilabel Prediction via Randomized Methods
- Fast Label Embeddings via Randomized Linear Algebra