1 paper
Jinguang Wang, Jingyu Wang, Haifeng Sun +6
Quantization has been widely used to compress and accelerate inference of large language models (LLMs). Existing methods focus on exploring the per-token dynamic calibration to ens…