3 papers
cs.LG2026
In-Context Reward Adaptation for Robust Preference Modeling
Zhenyu Sun, Zheng Xu, Ermin Wei
Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. However, human values are inherent…
cs.LG2026
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback
Qiang Liu, Adrienne Kline, Ermin Wei
Safe Reinforcement Learning from Human Feedback (Safe RLHF) has recently achieved empirical success in developing helpful and harmless large language models by decoupling human pre…
cs.LG2024
Debiasing Federated Learning with Correlated Client Participation
Zhenyu Sun, Ziyang Zhang, Zheng Xu +3
In cross-device federated learning (FL) with millions of mobile clients, only a small subset of clients participate in training in every communication round, and Federated Averagin…