Distributed Training and Optimization Of Neural Networks
arXiv:2012.01839 · doi:10.1142/9789811234033_0008
Abstract
Deep learning models are yielding increasingly better performances thanks to multiple factors. To be successful, model may have large number of parameters or complex architectures and be trained on large dataset. This leads to large requirements on computing resource and turn around time, even more so when hyper-parameter optimization is done (e.g search over model architectures). While this is a challenge that goes beyond particle physics, we review the various ways to do the necessary computations in parallel, and put it in the context of high energy physics.
20 pages, 4 figures, 2 tables, Submitted for review. To appear in "Artificial Intelligence for Particle Physics", World Scientific Publishing
References in corpus (11)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Practical Bayesian Optimization of Machine Learning Algorithms
- Population Based Training of Neural Networks
- An Empirical Model of Large-Batch Training
- A Living Review of Machine Learning for Particle Physics
- Graph Neural Networks for Particle Tracking and Reconstruction
- torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models
- Exascale Deep Learning for Scientific Inverse Problems
- High Resolution Medical Image Analysis with Spatial Partitioning
- An MPI-Based Python Framework for Distributed Training with Keras
- Randomized Automatic Differentiation