1 paper
Jing Liu, Toshiaki Koike-Akino, Ye Wang +2
To address the enormous size of Large Language Models (LLMs), model compression methods, such as quantization and pruning, are often deployed, especially on edge devices. In this w…