2 citations · 3 across the 4 of their papers we have counts for
1 paper · 2 filters
Wenhua Cheng, Weiwei Zhang, Heng Guo +2
Extremely low-bit quantization is critical for efficiently deploying Large Language Models (LLMs), yet it often leads to severe performance degradation at 2 bits and even at 4 bits…