4 papers
Compressing LLMs with MoP: Mixture of Pruners
Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias +7
The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, model pruning emerges as an effec…
Layer-wise LoRA fine-tuning: a similarity metric approach
Keith Ando Ogawa, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara +5
Pre-training Large Language Models (LLMs) on web-scale datasets becomes fundamental for advancing general-purpose AI. In contrast, enhancing their predictive performance on downstr…
Technical Report on Text Dataset Distillation
Keith Ando Ogawa, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara +6
In the vision domain, dataset distillation arises as a technique to condense a large dataset into a smaller synthetic one that exhibits a similar result in the training process. Wh…
Efficient LLMs with AMP: Attention Heads and MLP Pruning
Leandro Giusti Mugnaini, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara +5
Deep learning drives a new wave in computing systems and triggers the automation of increasingly complex problems. In particular, Large Language Models (LLMs) have significantly ad…