1 paper
Yipin Guo, Arun M George, Jie Fu +3
Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to ge…