activity
20232026
most citedTEQ: Trainable Equivalent Transformation for Quantization of LLMs

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2026

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs

Yu Luo, Bo Dong, Wenhua Cheng +1

Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional qua…

cs.CL2025

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs

Wenhua Cheng, Weiwei Zhang, Heng Guo +2

Extremely low-bit quantization is critical for efficiently deploying Large Language Models (LLMs), yet it often leads to severe performance degradation at 2 bits and even at 4 bits…

cs.CV2023

Effective Quantization for Diffusion Models on CPUs

Hanwen Chang, Haihao Shen, Yiyang Cai +7

Diffusion models have gained popularity for generating images from textual descriptions. Nonetheless, the substantial need for computational resources continues to present a notewo…

cs.CL20231 cited

TEQ: Trainable Equivalent Transformation for Quantization of LLMs

Wenhua Cheng, Yiyang Cai, Kaokao Lv +1

As large language models (LLMs) become more prevalent, there is a growing need for new and improved quantization methods that can meet the computationalast layer demands of these m…

cs.CL2023

Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Wenhua Cheng, Weiwei Zhang, Haihao Shen +4

Large Language Models (LLMs) have demonstrated exceptional proficiency in language-related tasks, but their deployment poses significant challenges due to substantial memory and st…