5 papers
From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
Elia Cunegatti, Marcus Vukojevic, Erik Nielsen +1
Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules. Existing replacement-…
Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization
Francesco Pio Monaco, Elia Cunegatti, Flavio Vella +1
Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance. While several compression approaches…
Hallucination as an Anomaly: Dynamic Intervention via Probabilistic Circuits
Erik Nielsen, Elia Cunegatti, Marcus Vukojevic +1
One of the most critical challenges in Large Language Models is their tendency to hallucinate, i.e., produce factually incorrect responses. Existing approaches show promising resul…
Zeroth-Order Adaptive Neuron Alignment Based Pruning without Re-Training
Elia Cunegatti, Leonardo Lucio Custode, Giovanni Iacca
Network pruning focuses on algorithms that aim to reduce a given model's computational cost by removing a subset of its parameters while having minimal impact on performance. Throu…
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
Fabrizio Sandri, Elia Cunegatti, Giovanni Iacca
We propose a novel Two-Stage framework for Structured Pruning (\textsc{2SSP}) for pruning Large Language Models (LLMs), which combines two different strategies of pruning, namely W…