3 papers
Compressing LLMs with MoP: Mixture of Pruners
Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias +7
The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, model pruning emerges as an effec…
Efficient LLMs with AMP: Attention Heads and MLP Pruning
Leandro Giusti Mugnaini, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara +5
Deep learning drives a new wave in computing systems and triggers the automation of increasingly complex problems. In particular, Large Language Models (LLMs) have significantly ad…
Layer Pruning with Consensus: A Triple-Win Solution
Leandro Giusti Mugnaini, Carolina Tavares Duarte, Anna H. Reali Costa +1
Layer pruning offers a promising alternative to standard structured pruning, effectively reducing computational costs, latency, and memory footprint. While notable layer-pruning ap…