3 papers
Compressing LLMs with MoP: Mixture of Pruners
Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias +7
The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, model pruning emerges as an effec…
Technical Report on Text Dataset Distillation
Keith Ando Ogawa, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara +6
In the vision domain, dataset distillation arises as a technique to condense a large dataset into a smaller synthetic one that exhibits a similar result in the training process. Wh…
Efficient LLMs with AMP: Attention Heads and MLP Pruning
Leandro Giusti Mugnaini, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara +5
Deep learning drives a new wave in computing systems and triggers the automation of increasingly complex problems. In particular, Large Language Models (LLMs) have significantly ad…