Distributed TensorFlow with MPI
arXiv:1603.02339
Abstract
Machine Learning and Data Mining (MLDM) algorithms are becoming increasingly important in analyzing large volume of data generated by simulations, experiments and mobile devices. With increasing data volume, distributed memory systems (such as tightly connected supercomputers or cloud computing systems) are becoming important in designing in-memory and massively parallel MLDM algorithms. Yet, the majority of open source MLDM software is limited to sequential execution with a few supporting multi-core/many-core execution. In this paper, we extend recently proposed Google TensorFlow for execution on large scale clusters using Message Passing Interface (MPI). Our approach requires minimal changes to the TensorFlow runtime -- making the proposed implementation generic and readily usable to increasingly large users of TensorFlow. We evaluate our implementation using an InfiniBand cluster and several well knowndatasets. Our evaluation indicates the efficiency of our proposed implementation.
6 pages; fixed significant typo
References in corpus (3)
Cited by in corpus (9)
- From Distributed Machine Learning to Federated Learning: A Survey
- Deep Learning in the Automotive Industry: Applications and Tools
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Communication optimization strategies for distributed deep neural network training: A survey
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Gear Training: A new way to implement high-performance model-parallel training
- STEP : A Distributed Multi-threading Framework Towards Efficient Data Analytics
- Cloud Collectives: Towards Cloud-aware Collectives forML Workloads with Rank Reordering
- How to Train your DNN: The Network Operator Edition