Experiments on Parallel Training of Deep Neural Network using Model Averaging
arXiv:1507.01239
Abstract
In this work we apply model averaging to parallel training of deep neural network (DNN). Parallelization is done in a model averaging manner. Data is partitioned and distributed to different nodes for local model updates, and model averaging across nodes is done every few minibatches. We use multiple GPUs for data parallelization, and Message Passing Interface (MPI) for communication between nodes, which allows us to perform model averaging frequently without losing much time on communication. We investigate the effectiveness of Natural Gradient Stochastic Gradient Descent (NG-SGD) and Restricted Boltzmann Machine (RBM) pretraining for parallel training in model-averaging framework, and explore the best setups in term of different learning rate schedules, averaging frequencies and minibatch sizes. It is shown that NG-SGD and RBM pretraining benefits parameter-averaging based model training. On the 300h Switchboard dataset, a 9.3 times speedup is achieved using 16 GPUs and 17 times speedup using 32 GPUs with limited decoding accuracy loss.
References in corpus (1)
Cited by in corpus (21)
- On the Convergence of Local Descent Methods in Federated Learning
- A Field Guide to Federated Optimization
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- Variance Reduced Local SGD with Lower Communication Complexity
- Split Learning for collaborative deep learning in healthcare
- Collaborative Deep Learning in Fixed Topology Networks
- Communication optimization strategies for distributed deep neural network training: A survey
- Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD
- Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks
- An Experimental Study of Data Heterogeneity in Federated Learning Methods for Medical Imaging
- Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks
- DBS: Dynamic Batch Size For Distributed Deep Neural Network Training
- Collaborative Deep Learning Across Multiple Data Centers
- Distributed Learning for Melanoma Classification using Personal Health Train
- A Hybrid-Order Distributed SGD Method for Non-Convex Optimization to Balance Communication Overhead, Computational Complexity, and Convergence Rate
- Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
- Federated Deep AUC Maximization for Heterogeneous Data with a Constant Communication Complexity
- SecEL: Privacy-Preserving, Verifiable and Fault-Tolerant Edge Learning for Autonomous Vehicles
- CrossoverScheduler: Overlapping Multiple Distributed Training Applications in a Crossover Manner
- Toward High-Throughput Artificial Intelligence-Based Segmentation in Oncological PET Imaging
- Trade-offs of Local SGD at Scale: An Empirical Study