1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CL2024
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
Wei Shen, Xiaoying Zhang, Yuanshun Yao +3
Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…
cs.LG2024★ 1 cited
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
Xiaoying Zhang, Jean-Francois Ton, Wei Shen +2
We introduce Adversarial Policy Optimization (AdvPO), a novel solution to the pervasive issue of reward over-optimization in Reinforcement Learning from Human Feedback (RLHF) for L…
cs.LG2023
Uncertainty-Aware Instance Reweighting for Off-Policy Learning
Xiaoying Zhang, Junpu Chen, Hongning Wang +4
Off-policy learning, referring to the procedure of policy optimization with access only to logged feedback data, has shown importance in various real-world applications, such as se…