8 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.DC2019★ 5 cited
HyPar-Flow: Exploiting MPI and Keras for Scalable Hybrid-Parallel DNN Training using TensorFlow
Ammar Ahmad Awan, Arpan Jain, Quentin Anthony +2
To reduce training time of large-scale DNNs, scientists have started to explore parallelization strategies like data-parallelism, model-parallelism, and hybrid-parallelism. While d…
cs.DC2017★ 8 cited
Optimized Broadcast for Deep Learning Workloads on Dense-GPU InfiniBand Clusters: MPI or NCCL?
Ammar Ahmad Awan, Ching-Hsiang Chu, Hari Subramoni +1
Dense Multi-GPU systems have recently gained a lot of attention in the HPC arena. Traditionally, MPI runtimes have been primarily designed for clusters with a large number of nodes…