1 paper · 1 filter
Xuan Ding, Rui Sun, Yunjian Zhang +6
The ever-increasing computational demands and deployment costs of large language models (LLMs) have spurred numerous compressing methods. Compared to quantization and unstructured…