activity
20212024
most citedOutlier Suppression: Pushing the Limit of Low-bit Transformer Language Models

28 citations · 39 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2024★ 5 cited

Selective Focus: Investigating Semantics Sensitivity in Post-training Quantization for Lane Detection

Yunqian Fan, Xiuying Wei, Ruihao Gong +4

Lane detection (LD) plays a crucial role in enhancing the L2+ capabilities of autonomous driving, capturing widespread attention. The Post-Processing Quantization (PTQ) could facil…

cs.CL2023★ 6 cited

Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Xiuying Wei, Yunchen Zhang, Yuhang Li +4

Post-training quantization~(PTQ) of transformer language models faces significant challenges due to the existence of detrimental outliers in activations. We observe that these outl…

cs.LG2022★ 28 cited

Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models

Xiuying Wei, Yunchen Zhang, Xiangguo Zhang +5

Transformer architecture has become the fundamental element of the widespread natural language processing~(NLP) models. With the trends of large NLP models, the increasing memory a…

cs.CV2021

Distribution-sensitive Information Retention for Accurate Binary Neural Network

Haotong Qin, Xiangguo Zhang, Ruihao Gong +3

Model binarization is an effective method of compressing neural networks and accelerating their inference process. However, a significant performance gap still exists between the 1…

cs.CV2021

Diverse Sample Generation: Pushing the Limit of Generative Data-free Quantization

Haotong Qin, Yifu Ding, Xiangguo Zhang +3

Generative data-free quantization emerges as a practical compression approach that quantizes deep neural networks to low bit-width without accessing the real data. This approach ge…

cs.CV2021

Diversifying Sample Generation for Accurate Data-Free Quantization

Xiangguo Zhang, Haotong Qin, Yifu Ding +6

Quantization has emerged as one of the most prevalent approaches to compress and accelerate neural networks. Recently, data-free quantization has been widely studied as a practical…