5 papers
A Single Stepsize Suffices for Unprojected Linear TD(0): Simultaneous Robust and Fast Rates via Polyak--Ruppert Averaging
Wei-Cheng Lee, Francesco Orabona
We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory. We provide high-probability guarantees for a plain unprojected TD(0) algorithm w…
Offline and Online KL-Regularized RLHF under Differential Privacy
Yulian Wu, Rushil Thareja, Praneeth Vepakomma +1
In this paper, we study the offline and online settings of reinforcement learning from human feedback (RLHF) with KL-regularization -- a widely used objective function in large lan…
SquarePO: Differentially Private and Robust -Preference Optimization in Offline Direct Alignment
Xingyu Zhou, Yulian Wu, Wenqian Weng +1
In this paper, we theoretically study the offline alignment of language models with human preference feedback, under both preference label corruption and privacy protections. To th…
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
Xingyu Zhou, Yulian Wu, Francesco Orabona
In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corru…
Optimal Regret of Bernoulli Bandits under Global Differential Privacy
Achraf Azize, Yulian Wu, Junya Honda +3
As sequential learning algorithms are increasingly applied to real life, ensuring data privacy while maintaining their utilities emerges as a timely question. In this context, regr…