5 papers
Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
Irina Proskurina, Guillaume Metzler, Antoine Gourru +1
Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show tha…
Fair-GPTQ: Bias-Aware Quantization for Large Language Models
Irina Proskurina, Guillaume Metzler, Julien Velcin
The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However…
Histoires Morales: A French Dataset for Assessing Moral Alignment
Thibaud Leteno, Irina Proskurina, Antoine Gourru +4
Aligning language models with human values is crucial, especially as they become more integrated into everyday life. While models are often adapted to user preferences, it is equal…
When Quantization Affects Confidence of Large Language Models?
Irina Proskurina, Luc Brun, Guillaume Metzler +1
Recent studies introduced effective compression techniques for Large Language Models (LLMs) via post-training quantization or low-bit weight representation. Although quantized weig…
Mini Minds: Exploring Bebeshka and Zlata Baby Models
Irina Proskurina, Guillaume Metzler, Julien Velcin
In this paper, we describe the University of Lyon 2 submission to the Strict-Small track of the BabyLM competition. The shared task is created with an emphasis on small-scale langu…