5 citations · 11 across the 4 of their papers we have counts for
6 papers
Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
Siyu Wang, Yi Rong, Shiqing Fan +6
The last decade has witnessed growth in the computational requirements for training deep neural networks. Current approaches (e.g., data/model parallelism, pipeline parallelism) pa…
DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging
Qinggang Zhou, Yawen Zhang, Pengcheng Li +4
The state-of-the-art deep learning algorithms rely on distributed training systems to tackle the increasing sizes of models and training data sets. Minibatch stochastic gradient de…
FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
Guoping Long, Jun Yang, Wei Lin
Performance optimization is the art of continuous seeking a harmonious mapping between the application domain and hardware. Recent years have witnessed a surge of deep learning (DL…
Characterizing Deep Learning Training Workloads on Alibaba-PAI
Mengdi Wang, Chen Meng, Guoping Long +4
Modern deep learning models have been exploited in various domains, including computer vision (CV), natural language processing (NLP), search and recommendation. In practical AI cl…
Graph-Adaptive Pruning for Efficient Inference of Convolutional Neural Networks
Mengdi Wang, Qing Zhang, Jun Yang +2
In this work, we propose a graph-adaptive pruning (GAP) method for efficient inference of convolutional neural networks (CNNs). In this method, the network is viewed as a computati…
FusionStitching: Deep Fusion and Code Generation for Tensorflow Computations on GPUs
Guoping Long, Jun Yang, Kai Zhu +1
In recent years, there is a surge on machine learning applications in industry. Many of them are based on popular AI frameworks like Tensorflow, Torch, Caffe, or MxNet, etc, and ar…