4 papers
The landscape of deterministic and stochastic optimal control problems: One-shot Optimization versus Dynamic Programming
Jihun Kim, Yuhao Ding, Yingjie Bi +1
Optimal control problems can be solved via a one-shot (single) optimization or a sequence of optimization using dynamic programming (DP). However, the computation of their global o…
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
Yuhao Ding, Junzi Zhang, Hyunin Lee +1
Entropy regularization is an efficient technique for encouraging exploration and preventing a premature convergence of (vanilla) policy gradient methods in reinforcement learning (…
A CMDP-within-online framework for Meta-Safe Reinforcement Learning
Vanshaj Khattar, Yuhao Ding, Bilgehan Sel +2
Meta-reinforcement learning has widely been used as a learning-to-learn framework to solve unseen tasks with limited experience. However, the aspect of constraint violations has no…
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
Donghao Ying, Mengzi Amy Guo, Hyunin Lee +3
We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measu…