A Survey on Distributed Machine Learning
arXiv:1912.09789 · doi:10.1145/3377454
Abstract
The demand for artificial intelligence has grown significantly over the last decade and this growth has been fueled by advances in machine learning techniques and the ability to leverage hardware acceleration. However, in order to increase the quality of predictions and render machine learning solutions feasible for more complex applications, a substantial amount of training data is required. Although small machine learning models can be trained with modest amounts of data, the input for training larger models such as neural networks grows exponentially with the number of parameters. Since the demand for processing training data has outpaced the increase in computation power of computing machinery, there is a need for distributing the machine learning workload across multiple machines, and turning the centralized into a distributed system. These distributed systems present new challenges, first and foremost the efficient parallelization of the training process and the creation of a coherent model. This article provides an extensive overview of the current state-of-the-art in the field by outlining the challenges and opportunities of distributed machine learning over conventional (centralized) machine learning, discussing the techniques used for distributed machine learning, and providing an overview of the systems that are available.
References in corpus (23)
- Improving neural networks by preventing co-adaptation of feature detectors
- Deep Learning with Differential Privacy
- Practical Bayesian Optimization of Machine Learning Algorithms
- End to End Learning for Self-Driving Cars
- Popular Ensemble Methods: An Empirical Study
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- Probabilistic Latent Semantic Analysis
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- How To Backdoor Federated Learning
- Revisiting Distributed Synchronous SGD
- Horovod: fast and easy distributed deep learning in TensorFlow
- A Reliable Effective Terascale Linear Learning System
- Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization
- Distributed Coordinate Descent Method for Learning with Big Data
- Gossip Learning with Linear Models on Fully Distributed Data
- Mini-Batch Primal and Dual Methods for SVMs
- Scaling Deep Learning on GPU and Knights Landing clusters
- BSP vs MapReduce
- Machine Learning and Cloud Computing: Survey of Distributed and SaaS Solutions
- Epiphany-V: A 1024 processor 64-bit RISC System-On-Chip
- No Peek: A Survey of private distributed deep learning
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
Cited by in corpus (25)
- Edge Learning for B5G Networks with Distributed Signal Processing: Semantic Communication, Edge Computing, and Wireless Sensing
- From Distributed Machine Learning to Federated Learning: A Survey
- Pervasive AI for IoT applications: A Survey on Resource-efficient Distributed Artificial Intelligence
- A Survey on Device Behavior Fingerprinting: Data Sources, Techniques, Application Scenarios, and Datasets
- Wireless Control for Smart Manufacturing: Recent Approaches and Open Challenges
- Responsible and Regulatory Conform Machine Learning for Medicine: A Survey of Challenges and Solutions
- Roadmap for Edge AI: A Dagstuhl Perspective
- Similarity-based Label Inference Attack against Training and Inference of Split Learning
- FedHe: Heterogeneous Models and Communication-Efficient Federated Learning
- Machine Learning Systems in the IoT: Trustworthiness Trade-offs for Edge Intelligence
- Communication-efficient Quantum Algorithm for Distributed Machine Learning
- Learning Forward Reuse Distance
- Parameter-Parallel Distributed Variational Quantum Algorithm
- LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
- Compression and Data Similarity: Combination of Two Techniques for Communication-Efficient Solving of Distributed Variational Inequalities
- HeterPS: Distributed Deep Learning With Reinforcement Learning Based Scheduling in Heterogeneous Environments
- HAMLET: A Hierarchical Agent-based Machine Learning Platform
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
- Compare Where It Matters: Using Layer-Wise Regularization To Improve Federated Learning on Heterogeneous Data
- Communication, Computing, Caching, and Sensing for Next Generation Aerial Delivery Networks
- On the Convergence of Quantized Parallel Restarted SGD for Central Server Free Distributed Training
- Machine Learning Systems for Intelligent Services in the IoT: A Survey
- Hyperdimensional Computing for Efficient Distributed Classification with Randomized Neural Networks
- Cost-efficient and Skew-aware Data Scheduling for Incremental Learning in 5G Network
- On the Convergence of Inexact Gradient Descent with Controlled Synchronization Steps