3 papers
cs.LG2026
DP-OPD: Differentially Private On-Policy Distillation for Language Models
Fatemeh Khadem, Sajad Mousavi, Yi Fang +1
Large language models (LLMs) are increasingly adapted to proprietary and domain-specific corpora that contain sensitive information, creating a tension between formal privacy guara…
cs.CR2025
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
Zhan Shi, Yefeng Yuan, Yuhong Liu +2
The performance of modern machine learning systems depends on access to large, high-quality datasets, often sourced from user-generated content or proprietary, domain-specific corp…
cs.CL2025
Measuring Large Language Models Capacity to Annotate Journalistic Sourcing
Subramaniam Vincent, Phoebe Wang, Zhan Shi +2
Since the launch of ChatGPT in late 2022, the capacities of Large Language Models and their evaluation have been in constant discussion and evaluation both in academic research and…