5 papers
Improved Bounds for Private and Robust Alignment
Wenqian Weng, Yi He, Xingyu Zhou
In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on the suboptimality gap in both offline and…
When Determinants Are Not Enough: Private Rare Switching
Xingyu Zhou
In this note, I would like to share a small research moment where Codex helped me find the right way to adapt rare switching to the private setting. The standard determinant-based…
Towards Differentially Private Reinforcement Learning with General Function Approximation
Yi He, Xingyu Zhou
We present the first theoretical guarantees for differentially private online reinforcement learning (RL) with general function approximation, extending beyond prior work restricte…
SquarePO: Differentially Private and Robust -Preference Optimization in Offline Direct Alignment
Xingyu Zhou, Yulian Wu, Wenqian Weng +1
In this paper, we theoretically study the offline alignment of language models with human preference feedback, under both preference label corruption and privacy protections. To th…
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
Xingyu Zhou, Yulian Wu, Francesco Orabona
In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corru…