99 citations · 242 across the 19 of their papers we have counts for
8 papers · 1 filter
Towards Scalable Distributed Training of Deep Learning on Public Cloud Clusters
Shaohuai Shi, Xianhao Zhou, Shutao Song +21
Distributed training techniques have been widely deployed in large-scale deep neural networks (DNNs) training on dense-GPU clusters. However, on public cloud clusters, due to the m…
Performance Characterization and Bottleneck Analysis of Hyperledger Fabric
Canhui Wang, Xiaowen Chu
Hyperledger Fabric is a popular open-source project for deploying permissioned blockchains. Many performance characteristics of the latest Hyperledger Fabric, such as performance c…
Efficient Sparse-Dense Matrix-Matrix Multiplication on GPUs Using the Customized Sparse Storage Format
Shaohuai Shi, Qiang Wang, Xiaowen Chu
Multiplication of a sparse matrix to a dense matrix (SpDM) is widely used in many areas like scientific computing and machine learning. However, existing works under-look the perfo…
A Quantitative Survey of Communication Optimizations in Distributed Deep Learning
Shaohuai Shi, Zhenheng Tang, Xiaowen Chu +3
Nowadays, large and complex deep learning (DL) models are increasingly trained in a distributed manner across multiple worker machines, in which extensive communications between wo…
FADNet: A Fast and Accurate Network for Disparity Estimation
Qiang Wang, Shaohuai Shi, Shizhen Zheng +2
Deep neural networks (DNNs) have achieved great success in the area of computer vision. The disparity estimation problem tends to be addressed by DNNs which achieve much better pre…
Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs
Qiang Wang, Shaohuai Shi, Canhui Wang +1
Distributed Deep Learning (DDL) has rapidly grown its popularity since it helps boost the training performance on high-performance GPU clusters. Efficient job scheduling is indispe…