1 paper
Junhan Kim, Gukryeol Lee, Seungwoo Son +2
Group-wise quantization is an effective strategy for mitigating accuracy degradation in low-bit quantization of large language models (LLMs). Among existing methods, GPTQ has been…