4 citations · 4 across the 3 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2023
SUBP: Soft Uniform Block Pruning for 1xN Sparse CNNs Multithreading Acceleration
Jingyang Xiang, Siqi Li, Jun Chen +4
The study of sparsity in Convolutional Neural Networks (CNNs) has become widespread to compress and accelerate models in environments with limited resources. By constraining N cons…
cs.LG2023
Unified Data-Free Compression: Pruning and Quantization without Fine-Tuning
Shipeng Bai, Jun Chen, Xintian Shen +2
Structured pruning and quantization are promising approaches for reducing the inference time and memory footprint of neural networks. However, most existing methods require the ori…
cs.LG2023
Data-Free Quantization via Mixed-Precision Compensation without Fine-Tuning
Jun Chen, Shipeng Bai, Tianxin Huang +3
Neural network quantization is a very promising solution in the field of model compression, but its resulting accuracy highly depends on a training/fine-tuning process and requires…