1 paper
Xin Nie, Liang Dong, Haicheng Zhang +2
Weight quantization effectively reduces memory consumption and enable the deployment of Large Language Models on edge devices, yet existing hardware-friendly methods often rely on…