1 paper
Haoqi Wang, Lorenz K. Mueller, Jiawei Zhuang +2
Low-bit quantization has been widely adopted to accelerate the inference of large language models (LLMs) by significantly reducing computational cost and memory usage. However, act…