4 citations · 8 across the 3 of their papers we have counts for
4 papers
EasyScale: Accuracy-consistent Elastic Training for Deep Learning
Mingzhen Li, Wencong Xiao, Biao Sun +9
Distributed synchronized GPU training is commonly used for deep learning. The resource constraint of using a fixed number of GPUs makes large-scale training jobs suffer from long q…
FamilySeer: Towards Optimized Tensor Codes by Exploiting Computation Subgraph Similarity
Shanjun Zhang, Mingzhen Li, Hailong Yang +3
Deploying various deep learning (DL) models efficiently has boosted the research on DL compilers. The difficulty of generating optimized tensor codes drives DL compiler to ask for…
Accelerating Sparse Approximate Matrix Multiplication on GPUs
Xiaoyan Liu, Yi Liu, Ming Dun +4
Although the matrix multiplication plays a vital role in computational linear algebra, there are few efficient solutions for matrix multiplication of the near-sparse matrices. The…
The Deep Learning Compiler: A Comprehensive Survey
Mingzhen Li, Yi Liu, Xiaoyan Liu +7
The difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community. Several DL compilers…