29 citations · 57 across the 9 of their papers we have counts for
12 papers
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs
Menglu Yu, Ye Tian, Bo Ji +3
Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL com…
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling
Menglu Yu, Chuan Wu, Bo Ji +1
In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scal…
DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
Shiqing Fan, Yi Rong, Chen Meng +10
It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an e…
Distributed Machine Learning through Heterogeneous Edge Systems
Hanpeng Hu, Dan Wang, Chuan Wu
Many emerging AI applications request distributed machine learning (ML) among edge systems (e.g., IoT devices and PCs at the edge of the Internet), where data cannot be uploaded to…
Characterizing Deep Learning Training Workloads on Alibaba-PAI
Mengdi Wang, Chen Meng, Guoping Long +4
Modern deep learning models have been exploited in various domains, including computer vision (CV), natural language processing (NLP), search and recommendation. In practical AI cl…
DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
Yanghua Peng, Yixin Bao, Yangrui Chen +3
More and more companies have deployed machine learning (ML) clusters, where deep learning (DL) models are trained for providing various AI-driven services. Efficient resource sched…