1 paper
Yongge Ma, Guoan Wang, Feiyu Wang +5
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized mod…