1 paper
Aoying Zheng, Anqi Du, Zizhuang Deng +3
Model quantization is a key technique for reducing storage and inference costs in large language model deployment. However, recent studies show that the discretization and rounding…