35 citations · 37 across the 2 of their papers we have counts for
4 papers
Dual-side Sparse Tensor Core
Yang Wang, Chen Zhang, Zhiqiang Xie +3
Leveraging sparsity in deep neural network (DNN) models is promising for accelerating model inference. Yet existing GPUs can only leverage the sparsity from weights but not activat…
LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
Yihuan Mao, Yujing Wang, Chufan Wu +6
BERT is a cutting-edge language representation model pre-trained by a large corpus, which achieves superior performances on various natural language understanding tasks. However, a…
Deeper Insights into Weight Sharing in Neural Architecture Search
Yuge Zhang, Zejun Lin, Junyang Jiang +5
With the success of deep neural networks, Neural Architecture Search (NAS) as a way of automatic model design has attracted wide attention. As training every child model from scrat…
Balanced Sparsity for Efficient DNN Inference on GPU
Zhuliang Yao, Shijie Cao, Wencong Xiao +2
In trained deep neural networks, unstructured pruning can reduce redundant weights to lower storage cost. However, it requires the customization of hardwares to speed up practical…