Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
arXiv:1802.09941
Abstract
Deep Neural Networks (DNNs) are becoming an important tool in modern computing applications. Accelerating their training is a major challenge and techniques range from distributed algorithms to low-level circuit design. In this survey, we describe the problem from a theoretical perspective, followed by approaches for its parallelization. We present trends in DNN architectures and the resulting implications on parallelization strategies. We then review and model the different types of concurrency in DNNs: from the single operator, through parallelism in network inference and training, to distributed deep learning. We discuss asynchronous stochastic optimization, distributed system architectures, communication schemes, and neural architecture search. Based on those approaches, we extrapolate potential directions for parallelism in deep learning.
References in corpus (41)
- Deep Learning in Neural Networks: An Overview
- Distilling the Knowledge in a Neural Network
- Practical Bayesian Optimization of Machine Learning Algorithms
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Neural Architecture Search with Reinforcement Learning
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Progressive Growing of GANs for Improved Quality, Stability, and Variation
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- DARTS: Differentiable Architecture Search
- Deep Learning with Limited Numerical Precision
- cuDNN: Efficient Primitives for Deep Learning
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- One weird trick for parallelizing convolutional neural networks
- Efficient Neural Architecture Search via Parameter Sharing
- Don't Decay the Learning Rate, Increase the Batch Size
- Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
- Large Batch Training of Convolutional Networks
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- One Model To Learn Them All
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Simple And Efficient Architecture Search for Convolutional Neural Networks
- Progressive Neural Architecture Search
- TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
- Parallel training of DNNs with Natural Gradient and Parameter Averaging
- Meta Learning Shared Hierarchies
- On the Origin of Deep Learning
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Scaling Deep Learning on GPU and Knights Landing clusters
- How to scale distributed deep learning?
- Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters
- ImageNet Training in Minutes
- Dual Path Networks
- Generative Adversarial Parallelization
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- Communication-Optimal Convolutional Neural Nets
- Ease.ml: Towards Multi-tenant Resource Sharing for Machine Learning Workloads
- SparCML: High-Performance Sparse Communication for Machine Learning
- On the Performance of Network Parallel Training in Artificial Neural Networks
Cited by in corpus (37)
- Split learning for health: Distributed deep learning without sharing raw patient data
- Database Meets Deep Learning: Challenges and Opportunities
- Think Locally, Act Globally: Federated Learning with Local and Global Representations
- Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
- Vector Symbolic Architectures as a Computing Framework for Emerging Hardware
- Communication optimization strategies for distributed deep neural network training: A survey
- Characterizing Deep-Learning I/O Workloads in TensorFlow
- Near-Optimal Sparse Allreduce for Distributed Deep Learning
- Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
- Augment your batch: better training with larger batches
- Flare: Flexible In-Network Allreduce
- Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation
- DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation
- Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Graph Processing on FPGAs: Taxonomy, Survey, Challenges
- Tighter Theory for Local SGD on Identical and Heterogeneous Data
- XPipe: Efficient Pipeline Model Parallelism for Multi-GPU DNN Training
- FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters
- Union: An Automatic Workload Manager for Accelerating Network Simulation
- Gradient Descent with Compressed Iterates
- SparCML: High-Performance Sparse Communication for Machine Learning
- An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks
- LayerPipe: Accelerating Deep Neural Network Training by Intra-Layer and Inter-Layer Gradient Pipelining and Multiprocessor Scheduling
- Efficient ConvNets for Analog Arrays
- Performance Analysis and Comparison of Distributed Machine Learning Systems
- Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
- Accelerating Neural Network Training with Distributed Asynchronous and Selective Optimization (DASO)
- Scavenger: A Cloud Service for Optimizing Cost and Performance of ML Training
- Speeding up Deep Learning with Transient Servers
- Characterizing and Modeling Distributed Training with Transient Cloud GPU Servers
- Stochastic Distributed Optimization for Machine Learning from Decentralized Features
- Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Pebbles, Graphs, and a Pinch of Combinatorics: Towards Tight I/O Lower Bounds for Statically Analyzable Programs
- Hydra: A System for Large Multi-Model Deep Learning
- A Random Gossip BMUF Process for Neural Language Modeling