1 paper
Yijia Zhang, Sicheng Zhang, Shijie Cao +4
Large language models (LLMs) show great performance in various tasks, but face deployment challenges from limited memory capacity and bandwidth. Low-bit weight quantization can sav…