5 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Yijia Zhang, Sicheng Zhang, Shijie Cao +4
Large language models (LLMs) show great performance in various tasks, but face deployment challenges from limited memory capacity and bandwidth. Low-bit weight quantization can sav…