2 papers
cs.LG2025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
Xiyue Peng, Hengquan Guo, Jiawei Zhang +4
Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training s…
cs.LG2024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning
Honghao Wei, Xiyue Peng, Arnob Ghosh +1
We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to…