papers

Publications (5)

cs.DC2022

Toward Efficient Online Scheduling for Distributed Machine Learning Systems

Menglu Yu, Jia Liu, Chuan Wu +2

Recent years have witnessed a rapid growth of distributed machine learning (ML) frameworks, which exploit the massive parallelism of computing clusters to expedite ML training. How…

cs.DC2022

GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs

Menglu Yu, Ye Tian, Bo Ji +3

Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL com…

cs.PF2026

Optimus: A Generic Operator-Level PyTorch Model Transformation Framework

Menglu Yu, Jiaqi Xu, Yuzhen Huang +19

In large-scale industrial applications, deep learning models that power recommendation and ranking have complex and diverse model architectures. These models are continuously devel…

cs.DC2021

A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling

Menglu Yu, Chuan Wu, Bo Ji +1

In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scal…

cs.DC2022

On Scheduling Ring-All-Reduce Learning Jobs in Multi-Tenant GPU Clusters with Communication Contention

Menglu Yu, Bo Ji, Hridesh Rajan +1

Powered by advances in deep learning (DL) techniques, machine learning and artificial intelligence have achieved astonishing successes. However, the rapidly growing needs for DL al…