1 paper · 1 filter
Pietro Tropeano, Maria Maistro, Tuukka Ruotsalo +1
Pruning Large Language Models (LLMs) reduces memory and inference costs by removing parts of the network, producing smaller models that retain most of their accuracy. As attention…