1 paper
He Li, Jianhang Hong, Yuanzhuo Wu +2
Model compression methods are used to reduce the computation and energy requirements for Large Language Models (LLMs). Quantization Aware Training (QAT), an effective model compres…