163 citations · 531 across the 27 of their papers we have counts for
3 papers · 1 filter
Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule Search
Zhiqi Lin, Youshan Miao, Guanbin Xu +4
Increasingly complex and diverse deep neural network (DNN) models necessitate distributing the execution across multiple devices for training and inference tasks, and also require…
SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction
Zhiqi Lin, Youshan Miao, Guodong Liu +10
With the growing model size, deep neural networks (DNN) are increasingly trained over massive GPU accelerators, which demands a proper parallelization plan that transforms a DNN mo…
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee +3
With widespread advances in machine learning, a number of large enterprises are beginning to incorporate machine learning models across a number of products. These models are typic…