45 citations · 58 across the 3 of their papers we have counts for
4 papers
dPRO: A Generic Profiling and Optimization System for Expediting Distributed DNN Training
Hanpeng Hu, Chenyu Jiang, Yuchen Zhong +5
Distributed training using multiple devices (e.g., GPUs) has been widely adopted for learning DNN models over large datasets. However, the performance of large-scale distributed tr…
DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
Yanghua Peng, Yixin Bao, Yangrui Chen +3
More and more companies have deployed machine learning (ML) clusters, where deep learning (DL) models are trained for providing various AI-driven services. Efficient resource sched…
Online Job Scheduling in Distributed Machine Learning Clusters
Yixin Bao, Yanghua Peng, Chuan Wu +1
Nowadays large-scale distributed machine learning systems have been deployed to support various analytics and intelligence services in IT firms. To train a large dataset and derive…
Dynamic Scaling of Virtualized, Distributed Service Chains: A Case Study of IMS
Jingpu Duan, Chuan Wu, Franck Le +2
The emerging paradigm of network function virtualization advocates deploying virtualized network functions (VNF) on standard virtualization platforms for significant cost reduction…