activity
20132022
most citedDAPPLE: A Pipelined Data Parallel Approach for Training Large Models

29 citations · 57 across the 9 of their papers we have counts for

collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2022

GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs

Menglu Yu, Ye Tian, Bo Ji +3

Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL com…

cs.DC2021

A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling

Menglu Yu, Chuan Wu, Bo Ji +1

In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scal…

cs.DC202029 cited

DAPPLE: A Pipelined Data Parallel Approach for Training Large Models

Shiqing Fan, Yi Rong, Chen Meng +10

It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an e…

cs.DC20194 cited

Distributed Machine Learning through Heterogeneous Edge Systems

Hanpeng Hu, Dan Wang, Chuan Wu

Many emerging AI applications request distributed machine learning (ML) among edge systems (e.g., IoT devices and PCs at the edge of the Internet), where data cannot be uploaded to…

cs.DC201813 cited

Online Job Scheduling in Distributed Machine Learning Clusters

Yixin Bao, Yanghua Peng, Chuan Wu +1

Nowadays large-scale distributed machine learning systems have been deployed to support various analytics and intelligence services in IT firms. To train a large dataset and derive…