1 paper
Yesheng Liang, Haisheng Chen, Zihan Zhang +2
Post-training quantization (PTQ) compresses the weights and activations of large language models (LLMs) into low-precision representations to reduce memory footprint and accelerate…