131 citations · 131 across the 2 of their papers we have counts for
3 papers
cs.DC2022
A Simulation Platform for Multi-tenant Machine Learning Services on Thousands of GPUs
Ruofan Liang, Bingsheng He, Shengen Yan +1
Multi-tenant machine learning services have become emerging data-intensive workloads in data centers with heavy usage of GPU resources. Due to the large scale, many tuning paramete…
cs.DC2021★ 131 cited
Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters
Qinghao Hu, Peng Sun, Shengen Yan +2
Modern GPU datacenters are critical for delivering Deep Learning (DL) models and services in both the research community and industry. When operating a datacenter, optimization of…
cs.DC2019
Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
Peng Sun, Wansen Feng, Ruobing Han +2
It is important to scale out deep neural network (DNN) training for reducing model training time. The high communication overhead is one of the major performance bottlenecks for di…