1 paper
Hongyaoxing Gul, Lijuan Hu, Shuzi Niu +1
Traditional post-training quantization (PTQ) is considered an effective approach to reduce model size and accelerate inference of large-scale language models (LLMs). However, exist…