5 papers
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
Chenhua Fan, Jiahui Zhu, Yuhang Zhang +1
Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies t…
Robust Peak-cost Constrained Reinforcement Learning
Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar +3
We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a tra…
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
Jiahui Zhu, Kihyun Yu, Dabeen Lee +2
Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn op…
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
Xiyue Peng, Hengquan Guo, Jiawei Zhang +4
Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training s…
Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis
Qining Zhang, Honghao Wei, Lei Ying
In this paper, we study reinforcement learning from human feedback (RLHF) under an episodic Markov decision process with a general trajectory-wise reward model. We developed a mode…