4 papers · 1 filter
From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
Elia Cunegatti, Marcus Vukojevic, Erik Nielsen +1
Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules. Existing replacement-…
Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization
Francesco Pio Monaco, Elia Cunegatti, Flavio Vella +1
Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance. While several compression approaches…
Hallucination as an Anomaly: Dynamic Intervention via Probabilistic Circuits
Erik Nielsen, Elia Cunegatti, Marcus Vukojevic +1
One of the most critical challenges in Large Language Models is their tendency to hallucinate, i.e., produce factually incorrect responses. Existing approaches show promising resul…
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
Fabrizio Sandri, Elia Cunegatti, Giovanni Iacca
We propose a novel Two-Stage framework for Structured Pruning (\textsc{2SSP}) for pruning Large Language Models (LLMs), which combines two different strategies of pruning, namely W…