3 papers
cs.CL2026
Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution
Dimitri Kachler, Damien Sileo, Pascal Denis
With the growth of LLMs' (Large Language Models) capabilities, there has been an increasing push to curate high quality datasets by filtering samples in the training data. In gener…
cs.LG2024
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
Gaurav Maheshwari, Aurélien Bellet, Pascal Denis +1
In this paper, we introduce a data augmentation approach specifically tailored to enhance intersectional fairness in classification tasks. Our method capitalizes on the hierarchica…
cs.CL2022
Fair NLP Models with Differentially Private Text Encoders
Gaurav Maheshwari, Pascal Denis, Mikaela Keller +1
Encoded text representations often capture sensitive attributes about individuals (e.g., race or gender), which raise privacy concerns and can make downstream models unfair to cert…