5 papers
The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
Irina Proskurina, Antoine Gourru, Julien Velcin
Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As synthetic content increasin…
Fair-GPTQ: Bias-Aware Quantization for Large Language Models
Irina Proskurina, Guillaume Metzler, Julien Velcin
The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However…
Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration
Sayan Kumar Chaki, Antoine Gourru, Julien Velcin
Fairness in language models is typically studied as a property of a single, centrally optimized model. As large language models become increasingly agentic, we propose that fairnes…
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
Irina Proskurina, Marc-Antoine Carpentier, Julien Velcin
Optimization of offensive content moderation models for different types of hateful messages is typically achieved through continued pre-training or fine-tuning on new hate speech b…
Histoires Morales: A French Dataset for Assessing Moral Alignment
Thibaud Leteno, Irina Proskurina, Antoine Gourru +4
Aligning language models with human values is crucial, especially as they become more integrated into everyday life. While models are often adapted to user preferences, it is equal…