3 papers
cs.CL2026
Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution
Dimitri Kachler, Damien Sileo, Pascal Denis
With the growth of LLMs' (Large Language Models) capabilities, there has been an increasing push to curate high quality datasets by filtering samples in the training data. In gener…
cs.CL2026
Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization
Gabriel Loiseau, Damien Sileo, Damien Riquet +2
Anonymizing textual documents is a highly context-sensitive problem: the appropriate balance between privacy protection and utility preservation varies with the data domain, privac…
cs.CL2026
Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
Gabriel Loiseau, Damien Sileo, Damien Riquet +2
Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving natural language processing. Recent work has shown that large language models (LLMs)…