4 papers · 1 filter
On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
Jiancong Xiao, Ziniu Li, Xingyu Xie +4
Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes. Ho…
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
Xiang Li, Garrett Wen, Weiqing He +3
Text watermarks in large language models (LLMs) are an increasingly important tool for detecting synthetic text and distinguishing human-written content from LLM-generated text. Wh…
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
Jiancong Xiao, Zhekun Shi, Kaizhao Liu +2
Despite its empirical success, Reinforcement Learning from Human Feedback (RLHF) has been shown to violate almost all the fundamental axioms in social choice theory -- such as majo…
Minimax Estimation for Personalized Federated Learning: An Alternative between FedAvg and Local Training?
Shuxiao Chen, Qinqing Zheng, Qi Long +1
A widely recognized difficulty in federated learning arises from the statistical heterogeneity among clients: local datasets often originate from distinct yet not entirely unrelate…