activity
20182023
most citedA Survey of Deep Learning Techniques for Neural Machine Translation

99 citations · 223 across the 15 of their papers we have counts for

collaborators
Showing cs.DCShow all

9 papers · 1 filter

cs.DC20215 cited

Energy-aware Task Scheduling with Deadline Constraint in DVFS-enabled Heterogeneous Clusters

Xinxin Mei, Qiang Wang, Xiaowen Chu +3

Energy conservation of large data centers for high-performance computing workloads, such as deep learning with big data, is of critical significance, where cutting down a few perce…

cs.DC2020

Towards Scalable Distributed Training of Deep Learning on Public Cloud Clusters

Shaohuai Shi, Xianhao Zhou, Shutao Song +21

Distributed training techniques have been widely deployed in large-scale deep neural networks (DNNs) training on dense-GPU clusters. However, on public cloud clusters, due to the m…

cs.DC2020

Performance Characterization and Bottleneck Analysis of Hyperledger Fabric

Canhui Wang, Xiaowen Chu

Hyperledger Fabric is a popular open-source project for deploying permissioned blockchains. Many performance characteristics of the latest Hyperledger Fabric, such as performance c…

cs.DC20203 cited

Efficient Sparse-Dense Matrix-Matrix Multiplication on GPUs Using the Customized Sparse Storage Format

Shaohuai Shi, Qiang Wang, Xiaowen Chu

Multiplication of a sparse matrix to a dense matrix (SpDM) is widely used in many areas like scientific computing and machine learning. However, existing works under-look the perfo…

cs.DC2020

A Quantitative Survey of Communication Optimizations in Distributed Deep Learning

Shaohuai Shi, Zhenheng Tang, Xiaowen Chu +3

Nowadays, large and complex deep learning (DL) models are increasingly trained in a distributed manner across multiple worker machines, in which extensive communications between wo…

cs.DC20203 cited

Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs

Qiang Wang, Shaohuai Shi, Canhui Wang +1

Distributed Deep Learning (DDL) has rapidly grown its popularity since it helps boost the training performance on high-performance GPU clusters. Efficient job scheduling is indispe…