activity
20132022
most citedDAPPLE: A Pipelined Data Parallel Approach for Training Large Models

29 citations · 57 across the 9 of their papers we have counts for

collaborators

12 papers

cs.DC2022

GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs

Menglu Yu, Ye Tian, Bo Ji +3

Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL com…

cs.DC2021

A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling

Menglu Yu, Chuan Wu, Bo Ji +1

In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scal…

cs.DC202029 cited

DAPPLE: A Pipelined Data Parallel Approach for Training Large Models

Shiqing Fan, Yi Rong, Chen Meng +10

It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an e…

cs.DC20194 cited

Distributed Machine Learning through Heterogeneous Edge Systems

Hanpeng Hu, Dan Wang, Chuan Wu

Many emerging AI applications request distributed machine learning (ML) among edge systems (e.g., IoT devices and PCs at the edge of the Internet), where data cannot be uploaded to…

cs.PF20195 cited

Characterizing Deep Learning Training Workloads on Alibaba-PAI

Mengdi Wang, Chen Meng, Guoping Long +4

Modern deep learning models have been exploited in various domains, including computer vision (CV), natural language processing (NLP), search and recommendation. In practical AI cl…

cs.LG2019

DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters

Yanghua Peng, Yixin Bao, Yangrui Chen +3

More and more companies have deployed machine learning (ML) clusters, where deep learning (DL) models are trained for providing various AI-driven services. Efficient resource sched…