1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Wenxiang Lin, Juntao Huang, Luhan Zhang +5
Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective for 4-bit activations and 8-b…