1 paper
Rayyan Abdalla, Amir Hussein, Min Wu +1
Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigid weight-saliency assumption…