2 papers
cs.LG2025
Policy Filtration for RLHF to Mitigate Noise in Reward Models
Chuheng Zhang, Wei Shen, Li Zhao +4
While direct policy optimization methods exist, pioneering LLMs are fine-tuned with reinforcement learning from human feedback (RLHF) to generate better responses under the supervi…
cs.LG2025
Reward-Driven Interaction: Enhancing Proactive Dialogue Agents through User Satisfaction Prediction
Wei Shen, Xiaonan He, Chuheng Zhang +3
Reward-driven proactive dialogue agents require precise estimation of user satisfaction as an intrinsic reward signal to determine optimal interaction strategies. Specifically, thi…