A Survey on Large-scale Machine Learning
arXiv:2008.03911
Abstract
Machine learning can provide deep insights into data, allowing machines to make high-quality predictions and having been widely used in real-world applications, such as text mining, visual classification, and recommender systems. However, most sophisticated machine learning approaches suffer from huge time costs when operating on large-scale data. This issue calls for the need of {Large-scale Machine Learning} (LML), which aims to learn patterns from big data with comparable performance efficiently. In this paper, we offer a systematic survey on existing LML methods to provide a blueprint for the future developments of this area. We first divide these LML methods according to the ways of improving the scalability: 1) model simplification on computational complexities, 2) optimization approximation on computational efficiency, and 3) computation parallelism on computational capabilities. Then we categorize the methods in each perspective according to their targeted scenarios and introduce representative methods in line with intrinsic strategies. Lastly, we analyze their limitations and discuss potential directions as well as open issues that are promising to address in the future.
References in corpus (44)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distributed Representations of Words and Phrases and their Compositionality
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- ADADELTA: An Adaptive Learning Rate Method
- An overview of gradient descent optimization algorithms
- Quantum Machine Learning
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Quantum support vector machine for big data classification
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- On the Convergence of Adam and Beyond
- HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
- A Survey on Multi-view Learning
- Towards Federated Learning at Scale: System Design
- QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Horovod: fast and easy distributed deep learning in TensorFlow
- Adversarially Robust Generalization Requires More Data
- Stochastic Gradient Hamiltonian Monte Carlo
- Deep Image: Scaling up Image Recognition
- Convergence Rates of Inexact Proximal-Gradient Methods for Convex Optimization
- Slow Learners are Fast
- Distributed Coordinate Descent Method for Learning with Big Data
- Asynchronous Stochastic Gradient Descent with Delay Compensation
- Equilibrated adaptive learning rates for non-convex optimization
- An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
- Communication Complexity of Distributed Convex Learning and Optimization
- Towards Optimal One Pass Large Scale Learning with Averaged Stochastic Gradient Descent
- To understand deep learning we need to understand kernel learning
- Variance Reduction in SGD by Distributed Importance Sampling
- The Optimal Sample Complexity of PAC Learning
- A Primer on Coordinate Descent Algorithms
- A Lower Bound for the Optimization of Finite Sums
- Optimal Rates for Random Fourier Features
- Error Compensated Quantized SGD and its Applications to Large-scale Distributed Optimization
- Diving into the shallows: a computational perspective on large-scale shallow learning
- Iterative MapReduce for Large Scale Machine Learning
- Optimal mini-batch and step sizes for SAGA
- Computing Web-scale Topic Models using an Asynchronous Parameter Server
- AdderNet: Do We Really Need Multiplications in Deep Learning?
- Improved large-scale graph learning through ridge spectral sparsification
- Fast Matrix Factorization with Non-Uniform Weights on Missing Data