activity
20242026
collaborators

7 papers

cs.CL2026

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question…

cs.CL2026

The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data

Irina Proskurina, Antoine Gourru, Julien Velcin

Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As synthetic content increasin…

cs.CL2026

Fair-GPTQ: Bias-Aware Quantization for Large Language Models

Irina Proskurina, Guillaume Metzler, Julien Velcin

The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However…

cs.CL2026

HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection

Irina Proskurina, Marc-Antoine Carpentier, Julien Velcin

Optimization of offensive content moderation models for different types of hateful messages is typically achieved through continued pre-training or fine-tuning on new hate speech b…

cs.CL2026

RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts

Darya Kharlamova, Irina Proskurina

Many errors in student essays can be explained by influence from the native language (L1). L1 interference refers to errors influenced by a speaker's first language, such as using…

cs.CL2025

Histoires Morales: A French Dataset for Assessing Moral Alignment

Thibaud Leteno, Irina Proskurina, Antoine Gourru +4

Aligning language models with human values is crucial, especially as they become more integrated into everyday life. While models are often adapted to user preferences, it is equal…