99 citations · 202 across the 8 of their papers we have counts for
10 papers
Efficient Sparse-Dense Matrix-Matrix Multiplication on GPUs Using the Customized Sparse Storage Format
Shaohuai Shi, Qiang Wang, Xiaowen Chu
Multiplication of a sparse matrix to a dense matrix (SpDM) is widely used in many areas like scientific computing and machine learning. However, existing works under-look the perfo…
FADNet: A Fast and Accurate Network for Disparity Estimation
Qiang Wang, Shaohuai Shi, Shizhen Zheng +2
Deep neural networks (DNNs) have achieved great success in the area of computer vision. The disparity estimation problem tends to be addressed by DNNs which achieve much better pre…
Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs
Qiang Wang, Shaohuai Shi, Canhui Wang +1
Distributed Deep Learning (DDL) has rapidly grown its popularity since it helps boost the training performance on high-performance GPU clusters. Efficient job scheduling is indispe…
Communication-Efficient Decentralized Learning with Sparsification and Adaptive Peer Selection
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu
Distributed learning techniques such as federated learning have enabled multiple workers to train machine learning models together to reduce the overall training time. However, cur…
A Survey of Deep Learning Techniques for Neural Machine Translation
Shuoheng Yang, Yuxin Wang, Xiaowen Chu
In recent years, natural language processing (NLP) has got great development with deep learning techniques. In the sub-field of machine translation, a new approach named Neural Mac…
Understanding Top-k Sparsification in Distributed Deep Learning
Shaohuai Shi, Xiaowen Chu, Ka Chun Cheung +1
Distributed stochastic gradient descent (SGD) algorithms are widely deployed in training large-scale deep learning models, while the communication overhead among workers becomes th…