Publications (5)
Toward Efficient Online Scheduling for Distributed Machine Learning Systems
Menglu Yu, Jia Liu, Chuan Wu +2
Recent years have witnessed a rapid growth of distributed machine learning (ML) frameworks, which exploit the massive parallelism of computing clusters to expedite ML training. How…
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs
Menglu Yu, Ye Tian, Bo Ji +3
Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL com…
Optimus: A Generic Operator-Level PyTorch Model Transformation Framework
Menglu Yu, Jiaqi Xu, Yuzhen Huang +19
In large-scale industrial applications, deep learning models that power recommendation and ranking have complex and diverse model architectures. These models are continuously devel…
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling
Menglu Yu, Chuan Wu, Bo Ji +1
In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scal…
On Scheduling Ring-All-Reduce Learning Jobs in Multi-Tenant GPU Clusters with Communication Contention
Menglu Yu, Bo Ji, Hridesh Rajan +1
Powered by advances in deep learning (DL) techniques, machine learning and artificial intelligence have achieved astonishing successes. However, the rapidly growing needs for DL al…