29 citations · 57 across the 9 of their papers we have counts for
5 papers · 1 filter
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs
Menglu Yu, Ye Tian, Bo Ji +3
Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL com…
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling
Menglu Yu, Chuan Wu, Bo Ji +1
In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scal…
DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
Shiqing Fan, Yi Rong, Chen Meng +10
It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an e…
Distributed Machine Learning through Heterogeneous Edge Systems
Hanpeng Hu, Dan Wang, Chuan Wu
Many emerging AI applications request distributed machine learning (ML) among edge systems (e.g., IoT devices and PCs at the edge of the Internet), where data cannot be uploaded to…
Online Job Scheduling in Distributed Machine Learning Clusters
Yixin Bao, Yanghua Peng, Chuan Wu +1
Nowadays large-scale distributed machine learning systems have been deployed to support various analytics and intelligence services in IT firms. To train a large dataset and derive…