1 paper · 1 filter
Yifan Tan, Haoze Wang, Chao Yan +1
Model quantization has become a crucial technique to address the issues of large memory consumption and long inference times associated with LLMs. Mixed-precision quantization, whi…