11 citations · 39 across the 5 of their papers we have counts for
7 papers
Parameter-Efficient Sparsity for Large Language Models Fine-Tuning
Yuchao Li, Fuli Luo, Chuanqi Tan +4
With the dramatically increased number of parameters in language models, sparsity methods have received ever-increasing research focus to compress and accelerate the models. While…
You Only Compress Once: Towards Effective and Elastic BERT Compression via Exploit-Explore Stochastic Nature Gradient
Shaokun Zhang, Xiawu Zheng, Chenyi Yang +7
Despite superior performance on various natural language processing tasks, pre-trained models such as BERT are challenged by deploying on resource-constraint devices. Most existing…
Towards Compact CNNs via Collaborative Compression
Yuchao Li, Shaohui Lin, Jianzhuang Liu +7
Channel pruning and tensor decomposition have received extensive attention in convolutional neural network compression. However, these two techniques are traditionally deployed in…
INT8 Winograd Acceleration for Conv1D Equipped ASR Models Deployed on Mobile Devices
Yiwu Yao, Yuchao Li, Chengyu Wang +8
The intensive computation of Automatic Speech Recognition (ASR) models obstructs them from being deployed on mobile devices. In this paper, we present a novel quantized Winograd op…
Interpretable Neural Network Decoupling
Yuchao Li, Rongrong Ji, Shaohui Lin +5
The remarkable performance of convolutional neural networks (CNNs) is entangled with their huge number of uninterpretable parameters, which has become the bottleneck limiting the e…
Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
Shaohui Lin, Rongrong Ji, Yuchao Li +2
The success of convolutional neural networks (CNNs) in computer vision applications has been accompanied by a significant increase of computation and memory costs, which prohibits…