Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning
Honghao Wei, Xiyue Peng, Arnob Ghosh +1
We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to…
cs.LG2024
Model-Free, Regret-Optimal Best Policy Identification in Online CMDPs
Zihan Zhou, Honghao Wei, Lei Ying
This paper considers the best policy identification (BPI) problem in online Constrained Markov Decision Processes (CMDPs). We are interested in algorithms that are model-free, have…