Exascale Deep Learning for Scientific Inverse Problems
arXiv:1909.11150
Abstract
We introduce novel communication strategies in synchronous distributed Deep Learning consisting of decentralized gradient reduction orchestration and computational graph-aware grouping of gradient tensors. These new techniques produce an optimal overlap between computation and communication and result in near-linear scaling (0.93) of distributed training up to 27,600 NVIDIA V100 GPUs on the Summit Supercomputer. We demonstrate our gradient reduction techniques in the context of training a Fully Convolutional Neural Network to approximate the solution of a longstanding scientific inverse problem in materials imaging. The efficient distributed training on a dataset size of 0.5 PB, produces a model capable of an atomically-accurate reconstruction of materials, and in the process reaching a peak performance of 2.15(4) EFLOPS.
13 pages, 9 figures. Under review by the Systems and Machine Learning (SysML) Conference (SysML '20)
References in corpus (5)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Phase recovery and holographic image reconstruction using deep learning in neural networks
- cuDNN: Efficient Primitives for Deep Learning
- Generating Long Sequences with Sparse Transformers
- Reconstruction of 3-D Atomic Distortions from Electron Microscopy with Deep Learning