Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
Chenhua Fan, Jiahui Zhu, Yuhang Zhang +1
Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies t…
cs.LG2026
Robust Peak-cost Constrained Reinforcement Learning
Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar +3
We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a tra…