1 paper · 1 filter
Yujia Tong, Yuxi Wang, Yunyang Wan +3
Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existing evaluations focusing almo…