1 paper
Tao Lu, Haoyu Wang, Zonghui Wang +3
With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can acce…