29 citations · 83 across the 13 of their papers we have counts for
6 papers · 1 filter
PICASSO: Unleashing the Potential of GPU-centric Training for Wide-and-deep Recommender Systems
Yuanxing Zhang, Langshi Chen, Siran Yang +12
The development of personalized recommendation has significantly improved the accuracy of information matching and the revenue of e-commerce platforms. Recently, it has 2 trends: 1…
Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
Siyu Wang, Yi Rong, Shiqing Fan +6
The last decade has witnessed growth in the computational requirements for training deep neural networks. Current approaches (e.g., data/model parallelism, pipeline parallelism) pa…
DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
Shiqing Fan, Yi Rong, Chen Meng +10
It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an e…
FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
Guoping Long, Jun Yang, Wei Lin
Performance optimization is the art of continuous seeking a harmonious mapping between the application domain and hardware. Recent years have witnessed a surge of deep learning (DL…
AliGraph: A Comprehensive Graph Neural Network Platform
Rong Zhu, Kun Zhao, Hongxia Yang +5
An increasing number of machine learning tasks require dealing with large graph datasets, which capture rich and complex relationship among potentially billions of elements. Graph…
FusionStitching: Deep Fusion and Code Generation for Tensorflow Computations on GPUs
Guoping Long, Jun Yang, Kai Zhu +1
In recent years, there is a surge on machine learning applications in industry. Many of them are based on popular AI frameworks like Tensorflow, Torch, Caffe, or MxNet, etc, and ar…