1 paper
Wenxiang Lin, Juntao Huang, Luhan Zhang +5
Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective for 4-bit activations and 8-b…