1 paper
Jiayi Chen, Jieqi Shi, Jing Huo +1
The rapid progress of Large Language Models (LLMs) has brought substantial computational and memory demands, spurring the adoption of low-bit quantization. While 8-bit and 4-bit fo…