26 citations · 78 across the 6 of their papers we have counts for
9 papers
Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning
Bingbing Li, Zhenglun Kong, Tianyun Zhang +4
Pre-trained large-scale language models have increasingly demonstrated high accuracy on many natural language processing (NLP) tasks. However, the limited weight storage and comput…
RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile Devices
Wei Niu, Mengshu Sun, Zhengang Li +7
Mobile devices are becoming an important carrier for deep learning tasks, as they are being equipped with powerful, high-end mobile CPUs and GPUs. However, it is still a challengin…
BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
Xiaolong Ma, Zhengang Li, Yifan Gong +8
Accelerating DNN execution on various resource-limited computing platforms has been a long-standing problem. Prior works utilize l1-based group lasso or dynamic regularization such…
RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech Recognition
Peiyan Dong, Siyue Wang, Wei Niu +8
Recurrent neural networks (RNNs) based automatic speech recognition has nowadays become prevalent on mobile devices such as smart phones. However, previous RNN compression techniqu…
SS-Auto: A Single-Shot, Automatic Structured Weight Pruning Framework of DNNs with Ultra-High Efficiency
Zhengang Li, Yifan Gong, Xiaolong Ma +6
Structured weight pruning is a representative model compression technique of DNNs for hardware efficiency and inference accelerations. Previous works in this area leave great space…
A SOT-MRAM-based Processing-In-Memory Engine for Highly Compressed DNN Implementation
Geng Yuan, Xiaolong Ma, Sheng Lin +2
The computing wall and data movement challenges of deep neural networks (DNNs) have exposed the limitations of conventional CMOS-based DNN accelerators. Furthermore, the deep struc…