21 citations · 35 across the 5 of their papers we have counts for
9 papers
Towards Efficient Post-training Quantization of Pre-trained Language Models
Haoli Bai, Lu Hou, Lifeng Shang +3
Network quantization has gained increasing attention with the rapid growth of large pre-trained language models~(PLMs). However, most existing quantization methods for PLMs follow…
Discrete Auto-regressive Variational Attention Models for Text Modeling
Xianghong Fang, Haoli Bai, Jian Li +3
Variational autoencoders (VAEs) have been widely applied for text modeling. In practice, however, they are troubled by two challenges: information underrepresentation and posterior…
BinaryBERT: Pushing the Limit of BERT Quantization
Haoli Bai, Wei Zhang, Lu Hou +6
The rapid development of large pre-trained language models has greatly increased the demand for model compression techniques, among which quantization is a popular solution. In thi…
Efficient Bitwidth Search for Practical Mixed Precision Neural Network
Yuhang Li, Wei Wang, Haoli Bai +3
Network quantization has rapidly become one of the most widely used methods to compress and accelerate deep neural networks. Recent efforts propose to quantize weights and activati…
RTN: Reparameterized Ternary Network
Yuhang Li, Xin Dong, Sai Qian Zhang +3
To deploy deep neural networks on resource-limited devices, quantization has been widely explored. In this work, we study the extremely low-bit networks which have tremendous speed…
Few Shot Network Compression via Cross Distillation
Haoli Bai, Jiaxiang Wu, Irwin King +1
Model compression has been widely adopted to obtain light-weighted deep neural networks. Most prevalent methods, however, require fine-tuning with sufficient training data to ensur…