1 paper
Xinzhe Zheng, Zhen-Qun Yang, Zishan Liu +4
Large Language Models (LLMs) deliver strong performance but are difficult to deploy under tight memory and compute constraints. Low-bit post-training quantization (PTQ) is a promis…