8 citations · 8 across the 1 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2018
Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation
Ammar Ahmad Awan, Jeroen Bedorf, Ching-Hsiang Chu +2
TensorFlow has been the most widely adopted Machine/Deep Learning framework. However, little exists in the literature that provides a thorough understanding of the capabilities whi…
cs.DC2017★ 8 cited
Optimized Broadcast for Deep Learning Workloads on Dense-GPU InfiniBand Clusters: MPI or NCCL?
Ammar Ahmad Awan, Ching-Hsiang Chu, Hari Subramoni +1
Dense Multi-GPU systems have recently gained a lot of attention in the HPC arena. Traditionally, MPI runtimes have been primarily designed for clusters with a large number of nodes…