1 paper
Fei Zuo, Xiaoyan Xi, Quanyi Zeng +2
Large language models are increasingly deployed on CPU-only platforms where memory bandwidth is the primary bottleneck for autoregressive generation. Weight quantization to four bi…