1 paper
Yuanhong He, Peiyu Niu, Jun Chen +2
As Large Language Models (LLMs) scale, weight-only quantization (W4A16: 4-bit weights, 16-bit activations) becomes critical for reducing memory footprint with minimal accuracy loss…