4 papers · 1 filter
In-Context Reward Adaptation for Robust Preference Modeling
Zhenyu Sun, Zheng Xu, Ermin Wei
Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. However, human values are inherent…
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback
Qiang Liu, Adrienne Kline, Ermin Wei
Safe Reinforcement Learning from Human Feedback (Safe RLHF) has recently achieved empirical success in developing helpful and harmless large language models by decoupling human pre…
Debiasing Federated Learning with Correlated Client Participation
Zhenyu Sun, Ziyang Zhang, Zheng Xu +3
In cross-device federated learning (FL) with millions of mobile clients, only a small subset of clients participate in training in every communication round, and Federated Averagin…
A Stochastic Quasi-Newton Method for Non-convex Optimization with Non-uniform Smoothness
Zhenyu Sun, Ermin Wei
Classical convergence analyses for optimization algorithms rely on the widely-adopted uniform smoothness assumption. However, recent experimental studies have demonstrated that man…