activity
20242026
collaborators

6 papers

cs.LG2026

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

Chenhua Fan, Jiahui Zhu, Yuhang Zhang +1

Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies t…

cs.LG2026

Robust Peak-cost Constrained Reinforcement Learning

Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar +3

We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a tra…

cs.LG2025

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

Jiahui Zhu, Kihyun Yu, Dabeen Lee +2

Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn op…

cs.LG2025

Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Xiyue Peng, Hengquan Guo, Jiawei Zhang +4

Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training s…

cs.LG2025

Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis

Qining Zhang, Honghao Wei, Lei Ying

In this paper, we study reinforcement learning from human feedback (RLHF) under an episodic Markov decision process with a general trajectory-wise reward model. We developed a mode…

cs.LG2024

Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning

Honghao Wei, Xiyue Peng, Arnob Ghosh +1

We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to…