4 papers
Algorithm to Compilation Co-design: An Integrated View of Neural Network Sparsity
Fu-Ming Guo, Austin Huang
Reducing computation cost, inference latency, and memory footprint of neural networks are frequently cited as research motivations for pruning and sparsity. However, operationalizi…
An Image Enhancing Pattern-based Sparsity for Real-time Inference on Mobile Devices
Xiaolong Ma, Wei Niu, Tianyun Zhang +8
Weight pruning has been widely acknowledged as a straightforward and effective method to eliminate redundancy in Deep Neural Networks (DNN), thereby achieving acceleration on vario…
Reweighted Proximal Pruning for Large-Scale Language Representation
Fu-Ming Guo, Sijia Liu, Finlay S. Mungall +2
Recently, pre-trained language representation flourishes as the mainstay of the natural language understanding community, e.g., BERT. These pre-trained language representations can…
PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices
Xiaolong Ma, Fu-Ming Guo, Wei Niu +5
Model compression techniques on Deep Neural Network (DNN) have been widely acknowledged as an effective way to achieve acceleration on a variety of platforms, and DNN weight prunin…