14 citations · 45 across the 8 of their papers we have counts for
6 papers · 1 filter
Energy-aware Task Scheduling with Deadline Constraint in DVFS-enabled Heterogeneous Clusters
Xinxin Mei, Qiang Wang, Xiaowen Chu +3
Energy conservation of large data centers for high-performance computing workloads, such as deep learning with big data, is of critical significance, where cutting down a few perce…
Efficient Sparse-Dense Matrix-Matrix Multiplication on GPUs Using the Customized Sparse Storage Format
Shaohuai Shi, Qiang Wang, Xiaowen Chu
Multiplication of a sparse matrix to a dense matrix (SpDM) is widely used in many areas like scientific computing and machine learning. However, existing works under-look the perfo…
Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs
Qiang Wang, Shaohuai Shi, Canhui Wang +1
Distributed Deep Learning (DDL) has rapidly grown its popularity since it helps boost the training performance on high-performance GPU clusters. Efficient job scheduling is indispe…
Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training
Yuxin Wang, Qiang Wang, Shaohuai Shi +4
Deep learning has become widely used in complex AI applications. Yet, training a deep neural network (DNNs) model requires a considerable amount of calculations, long running time,…
A Distributed Synchronous SGD Algorithm with Global Top- Sparsification for Low Bandwidth Networks
Shaohuai Shi, Qiang Wang, Kaiyong Zhao +4
Distributed synchronous stochastic gradient descent (S-SGD) has been widely used in training large-scale deep neural networks (DNNs), but it typically requires very high communicat…
A DAG Model of Synchronous Stochastic Gradient Descent in Distributed Deep Learning
Shaohuai Shi, Qiang Wang, Xiaowen Chu +1
With huge amounts of training data, deep learning has made great breakthroughs in many artificial intelligence (AI) applications. However, such large-scale data sets present comput…