1 paper
Xin Yuan, Siqi Li, Jiateng Wei +7
Pruning is an effective method for compressing Large Language Models, but finding an optimal, non-uniform layer-wise sparsity allocation remains a key challenge. While heuristic me…