1 paper · 1 filter
Yipin Guo, Arun M George, Jie Fu +3
Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to ge…