4 papers · 1 filter
Robust Peak-cost Constrained Reinforcement Learning
Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar +3
We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a tra…
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
Sourav Ganguly, Kartik Pandit, Arnob Ghosh
Real-world decision-making systems operate in environments where state transitions depend not only on the agent's actions, but also on \textbf{exogenous factors outside its control…
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
Amirhossein Roknilamouki, Arnob Ghosh, Eylem Ekici +1
While offline reinforcement learning provides reliable policies for real-world deployment, its inherent pessimism severely restricts an agent's ability to explore and collect novel…
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning
Honghao Wei, Xiyue Peng, Arnob Ghosh +1
We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to…